## Table: Performance Comparison of AI Models on Math and STEM Benchmarks
### Overview
This image presents a performance comparison table evaluating eight different AI models across seven distinct mathematical and STEM-related benchmarks, plus an "Average" row. The table highlights the performance of the "AceMath" and "Qwen2.5-Math" model families within a green rounded rectangle, suggesting these are the primary subjects of the analysis. The table includes standard performance metrics and, for the highlighted models, additional "rm@8" (likely Reward Model at 8 samples) metrics.
### Components/Axes
**Rows (Benchmarks):**
* **GSM8K:** Grade school math
* **MATH:** High school math competition
* **Minerva Math:** Undergraduate-level quantitative reasoning
* **Gaokao 2023 English:** College-entry math exam
* **Olympiad Bench:** Olympiad-level math reasoning
* **College Math:** College-level mathematics
* **MMLU STEM:** Undergraduate-level STEM knowledge
* **Average:** The mean performance across the listed benchmarks
**Columns (Models):**
* **Highlighted Group (Green Box):**
* AceMath 72B-Instruct
* AceMath 7B-Instruct
* AceMath 1.5B-Instruct
* Qwen2.5-Math 72B-Instruct
* Qwen2.5-Math 7B-Instruct
* **Comparison Group:**
* LLaMa3.1 405B-Instruct
* GPT-4o* (2024-08-06)
* Claude 3.5 Sonnet (2024-10-22)
**Visual Elements:**
* **Green Rounded Rectangle:** Encloses the first five columns (AceMath and Qwen2.5-Math variants).
* **Green Text:** Indicates the highest score achieved in that specific row.
* **Footnote:** Located at the bottom, stating: "* OpenAI's o1 model family is excluded in this table due to dependency on extensive pre-response computation mechanism."
### Detailed Analysis
The table provides two values for the models inside the green box: a primary score and a secondary score labeled "(rm@8)". The models outside the box only have a single score.
| Benchmark | AceMath 72B | AceMath 7B | AceMath 1.5B | Qwen2.5-Math 72B | Qwen2.5-Math 7B | LLaMa3.1 405B | GPT-4o* | Claude 3.5 |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **GSM8K** | 96.4 / 97.1 | 93.7 / 96.4 | 87.0 / 93.9 | 95.9 / 96.4 | 95.2 / 97.9 | 96.8 | 92.9 | 96.4 |
| **MATH** | 86.1 / 89.4 | 83.1 / 87.8 | 76.8 / 84.9 | 85.9 / 89.8 | 83.6 / 88.5 | 73.8 | 81.1 | 78.3 |
| **Minerva** | 57.0 / 59.9 | 51.1 / 55.2 | 41.5 / 49.3 | 44.1 / 47.4 | 37.1 / 42.6 | 54.0 | 50.7 | 48.2 |
| **Gaokao** | 72.2 / 76.1 | 68.1 / 76.1 | 64.4 / 71.4 | 71.9 / 76.9 | 66.8 / 75.1 | 62.1 | 67.5 | 64.9 |
| **Olympiad** | 48.4 / 52.0 | 42.2 / 50.2 | 33.8 / 46.2 | 49.0 / 54.5 | 41.6 / 49.9 | 34.8 | 43.3 | 37.9 |
| **College** | 57.3 / 59.6 | 56.6 / 59.5 | 54.4 / 58.3 | 49.5 / 50.6 | 46.8 / 49.6 | 49.3 | 48.5 | 48.5 |
| **MMLU STEM**| 85.4 / 89.4 | 75.3 / 86.0 | 62.0 / 81.6 | 80.8 / 80.1 | 71.9 / 78.7 | 83.1 | 87.9 | 85.1 |
| **Average** | 71.8 / 74.8 | 67.2 / 73.0 | 60.0 / 69.4 | 68.2 / 70.8 | 62.9 / 68.9 | 64.8 | 67.4 | 65.6 |
*(Note: Values are presented as Primary Score / (rm@8) Score)*
### Key Observations
* **Performance Dominance:** The AceMath 72B-Instruct model consistently achieves the highest scores in the majority of categories, often outperforming the comparison models (LLaMa, GPT-4o, Claude 3.5).
* **The "rm@8" Effect:** The "(rm@8)" values are consistently higher than the primary scores for all models within the green box. This indicates that the "rm@8" metric (likely representing a re-ranking or best-of-N sampling strategy) significantly boosts the reported performance.
* **Outliers:**
* **LLaMa3.1 405B** achieves the highest score in **GSM8K** (96.8).
* **GPT-4o*** achieves the highest score in **MMLU STEM** (87.9).
* **AceMath 72B** dominates the **Minerva Math**, **Gaokao**, **Olympiad Bench**, and **College Math** categories.
### Interpretation
This table is a promotional or technical comparison document intended to demonstrate the high efficacy of the AceMath model family, particularly in specialized mathematical reasoning tasks.
* **Strategic Framing:** By grouping AceMath and Qwen2.5-Math together in a green box, the document implies a lineage or shared architecture, positioning AceMath as the superior iteration.
* **Metric Selection:** The inclusion of the "(rm@8)" data is a deliberate choice to show "potential" performance. By highlighting these numbers, the authors demonstrate that their models have high ceiling performance when given multiple attempts or re-ranking capabilities.
* **Contextual Exclusion:** The footnote regarding OpenAI's o1 model is a critical piece of information. It acknowledges the existence of a superior or differently-architected competitor (o1) and preemptively explains its absence, thereby protecting the validity of the AceMath results against the "o1" benchmark.
* **Conclusion:** The data suggests that while general-purpose models like GPT-4o and Claude 3.5 remain highly competitive in broad STEM knowledge (MMLU), the AceMath models are specifically fine-tuned to excel in rigorous, multi-step mathematical reasoning benchmarks (Olympiad, College Math, Minerva).