## Line Chart: D8 Per-Layer Q-Projection Stable Rank
### Overview
The chart visualizes the evolution of "Stable Rank" for eight layers (Layer 0–7) of a D8 model during training. Stable Rank is plotted against training steps (0–10,000). All layers exhibit a sharp initial decline in stable rank, followed by stabilization. Layer 7 (gray diamonds) consistently shows the highest stable rank, while Layers 0 and 4 (blue and light blue) achieve the lowest.
### Components/Axes
- **X-axis**: Training Step (0–10,000, linear scale)
- **Y-axis**: Stable Rank (0–140, linear scale)
- **Legend**:
- Layer 0: Solid blue circles
- Layer 1: Dashed orange squares
- Layer 2: Dotted green triangles
- Layer 3: Dotted pink diamonds
- Layer 4: Dashed light blue triangles
- Layer 5: Dashed yellow squares
- Layer 6: Dotted brown triangles
- Layer 7: Dotted gray diamonds
- **Legend Position**: Right-aligned, outside the plot area.
### Detailed Analysis
1. **Initial Decline (0–1,000 steps)**:
- All layers start near **130–140 stable rank** at step 0.
- Sharp drop to **~20–30** by step 1,000. Layer 7 (gray) declines most steeply, reaching ~40 by step 1,000.
- Layer 0 (blue) and Layer 4 (light blue) drop fastest, reaching ~15 by step 1,000.
2. **Stabilization (1,000–10,000 steps)**:
- All lines plateau with minimal fluctuation.
- Layer 7 (gray) stabilizes at **~20–25**, remaining the highest.
- Layers 0 and 4 stabilize at **~10–15**, the lowest.
- Layers 1–6 cluster between **15–25**, with Layer 1 (orange) and Layer 6 (brown) showing slight variability.
### Key Observations
- **Rapid Initial Drop**: All layers reduce stable rank by ~90% within the first 1,000 steps.
- **Layer 7 Anomaly**: Consistently highest stable rank post-stabilization (~20–25 vs. ~10–15 for others).
- **Layer 0/4 Performance**: Outperform other layers, achieving the lowest stable ranks.
- **Minor Fluctuations**: Layers 1–6 show small oscillations but remain within tight bounds post-1,000 steps.
### Interpretation
The chart demonstrates that the D8 model's Q-projections stabilize rapidly during early training, with most layers converging to low stable ranks by step 1,000. Layer 7's persistently higher stable rank suggests it may be less critical or more variable in the model's architecture. Layers 0 and 4, with the lowest stable ranks, likely play a more pivotal role in the model's stability. The uniformity of stabilization across layers implies robust training dynamics, though Layer 7's divergence warrants further investigation into its functional role or initialization.