Comparative Literature Review & Benchmark Synthesis
Synthesizes recent academic papers or model reports into a comparative matrix covering architecture, data, and benchmarks.
Act as a Research Scientist in Machine Learning. Analyze the provided excerpts from recent frontier model evaluation papers. Synthesize into an empirical comparative review: 1. Methodology Taxonomy: Classify each paper's primary evaluation mechanism (RLHF, DPO, CoT verification, synthetic data curation). 2. Benchmark Contrast: Build a markdown comparison table mapping common metrics (MMLU-Pro, MATH-500, HumanEval, SWE-bench). 3. Divergences & Confounds: Highlight discrepancies in test setups, prompt formatting, or test-time compute scaling. 4. Open Problems: Identify 2 unresolved limitations common across all reviewed approaches. Paper Excerpts: [INSERT EXCERPTS HERE]