## Line Graphs: Model Performance Across Training Steps
### Overview
The image contains four line graphs comparing the performance of two models (GSM8K and MMLU-Pro) under original and optimized configurations. Each graph tracks scores across three dataset sizes (0.96B, 2.07B, 4.14B) over 140K training steps. Data points are color-coded (red, green, blue) and marked with distinct symbols (circle, triangle, plus).
---
### Components/Axes
- **X-axis**: Training steps (K), ranging from 0 to 140K in increments of 20K.
- **Y-axis**: Score (numeric values, varying by subplot).
- **Legends**: Positioned on the right of each subplot, mapping:
- Red circles: 0.96B dataset
- Green triangles: 2.07B dataset
- Blue pluses: 4.14B dataset
- **Subplot Titles**:
1. GSM8K-Original
2. GSM8K-Optimized
3. MMLU-Pro-Original
4. MMLU-Pro-Optimized
---
### Detailed Analysis
#### GSM8K-Original
- **Red (0.96B)**: Starts at ~0, fluctuates between 0–10, peaks at ~8 at 100K, ends at ~5.
- **Green (2.07B)**: Starts at ~2, fluctuates between 2–20, peaks at ~25 at 120K, ends at ~22.
- **Blue (4.14B)**: Starts at ~5, fluctuates between 5–30, peaks at ~30 at 100K, ends at ~25.
#### GSM8K-Optimized
- **Red (0.96B)**: Starts at ~2, steadily increases to ~12 by 140K.
- **Green (2.07B)**: Starts at ~5, rises to ~25 by 140K with minor fluctuations.
- **Blue (4.14B)**: Starts at ~12, peaks at ~30 at 100K, drops slightly to ~28 at 140K.
#### MMLU-Pro-Original
- **Red (0.96B)**: Starts at ~11, fluctuates between 10–12, ends at ~11.
- **Green (2.07B)**: Starts at ~11, rises to ~14 at 100K, drops to ~12 at 140K.
- **Blue (4.14B)**: Starts at ~11, sharply rises to ~16 at 100K, drops to ~15 at 140K.
#### MMLU-Pro-Optimized
- **Red (0.96B)**: Starts at ~9, steadily increases to ~13 by 140K.
- **Green (2.07B)**: Starts at ~11, rises to ~15 at 100K, stabilizes at ~14 at 140K.
- **Blue (4.14B)**: Starts at ~11, sharply rises to ~16 at 100K, peaks at ~17 at 140K.
---
### Key Observations
1. **Optimized Models Outperform Originals**:
- GSM8K-Optimized scores are consistently higher than GSM8K-Original.
- MMLU-Pro-Optimized shows smoother growth compared to the volatile original.
2. **Dataset Size Correlation**:
- Larger datasets (4.14B) generally achieve higher scores, especially in optimized models.
- 0.96B datasets underperform across all configurations.
3. **Training Dynamics**:
- Original models exhibit erratic fluctuations, while optimized models show steadier improvement.
- MMLU-Pro-Optimized demonstrates the most consistent upward trend.
---
### Interpretation
The data suggests that model optimization significantly improves performance stability and final scores. Larger datasets (4.14B) yield higher scores, indicating scalability benefits. However, the 0.96B dataset underperforms, possibly due to insufficient capacity. The MMLU-Pro-Optimized model’s steady growth implies effective learning dynamics, whereas the GSM8K-Optimized model’s peak at 100K followed by a slight decline may signal overfitting or resource constraints. These trends highlight the trade-offs between dataset size, model optimization, and training efficiency.