## Radar Chart: Capabilities on the HEAR Benchmark
### Overview
The chart compares the performance of five audio processing models (Dasheng-1.2B, CED-Base, Whisper-base, AudioMAE, Wav2Vec2) across 15 HEAR benchmark tasks. Each model is represented by a distinct colored line, with performance scores plotted on a radial scale from 0 to 100. The chart emphasizes task-specific capabilities through overlapping performance distributions.
### Components/Axes
- **Title**: "Capabilities on the HEAR Benchmark" (top-center)
- **Legend**: Right-aligned, mapping colors to models:
- Blue: Dasheng-1.2B
- Orange: CED-Base
- Green: Whisper-base
- Red: AudioMAE
- Purple: Wav2Vec2
- **Axes**: 15 radial categories (clockwise from top):
1. Beijing Opera Percussion (96.6)
2. CREMA-D (81.6)
3. DCASE16 (94.2)
4. ESC-50 (65.5)
5. FSD50K (88.0)
6. GTZAN Genre (88.8)
7. GTZAN Music Speech (97.7)
8. LibriCount (79.6)
9. Maestro 5hr (43.3)
10. Mridangam Stroke (97.4)
11. Mridangam Tonic (96.5)
12. NSynth Pitch 50hr (85.6)
13. NSynth Pitch 5hr (74.4)
14. Speech Commands 5hr (97.1)
15. Speech Commands Full (97.9)
16. Vocal Imitations (92.7)
17. VoxLingua107 Top10 (88.7)
### Detailed Analysis
1. **Dasheng-1.2B (Blue)**:
- Dominates Speech Commands (97.9, 97.1) and GTZAN Music Speech (97.7)
- Weakest in Maestro 5hr (43.3) and ESC-50 (65.5)
- Consistent mid-to-high performance across most tasks
2. **CED-Base (Orange)**:
- Peaks in Vocal Imitations (97.9) and Mridangam Stroke (97.4)
- Struggles with Maestro 5hr (43.3) and ESC-50 (65.5)
- Notable dip in Speech Commands Full (92.7)
3. **Whisper-base (Green)**:
- Most consistent performer (range: 81.6–97.4)
- Strong in VoxLingua107 Top10 (88.7) and Mridangam Tonic (96.5)
- Weakest in ESC-50 (65.5) and Maestro 5hr (43.3)
4. **AudioMAE (Red)**:
- Excels in CREMA-D (81.6) and DCASE16 (94.2)
- Poor performance in Maestro 5hr (43.3) and ESC-50 (65.5)
- Notable strength in Speech Commands 5hr (97.1)
5. **Wav2Vec2 (Purple)**:
- Balanced performance (range: 79.6–99.1)
- Best in GTZAN Genre (88.8) and Mridangam Stroke (97.4)
- Weakest in Maestro 5hr (43.3) and ESC-50 (65.5)
### Key Observations
- **Speech Command Dominance**: All models score ≥92.7 on Speech Commands tasks, suggesting this is a well-addressed benchmark category.
- **Cultural Task Variance**: Models show divergent strengths in culturally specific tasks:
- CED-Base leads in Vocal Imitations (97.9)
- Whisper-base dominates Mridangam tasks (96.5–97.4)
- AudioMAE performs best on CREMA-D (81.6)
- **Maestro 5hr Weakness**: All models score ≤43.3 on Maestro 5hr, indicating a potential benchmark limitation or task complexity.
- **ESC-50 Bottleneck**: All models score ≤65.5 on ESC-50, highlighting a universal challenge in environmental sound classification.
### Interpretation
The chart reveals task-specific model specializations:
- **Dasheng-1.2B** excels in music-related tasks (GTZAN Music Speech: 97.7) and general speech commands.
- **CED-Base** specializes in vocal synthesis (Vocal Imitations: 97.9) and percussion tasks.
- **Whisper-base** offers reliability across diverse tasks but lacks peak performance.
- **AudioMAE** shows strength in acoustic scene understanding (CREMA-D: 81.6).
- **Wav2Vec2** demonstrates balanced capabilities with notable cultural task performance (Mridangam Stroke: 97.4).
The Maestro 5hr and ESC-50 weaknesses across all models suggest potential limitations in the HEAR benchmark's representation of complex musical analysis and environmental sound recognition. The cultural task variations (e.g., Mridangam vs. Beijing Opera) highlight the importance of model specialization for regional audio processing applications.