## Bar Chart: Peak Memory Usage vs Sequence Length
### Overview
The chart compares peak memory usage (in GB) across three compression strategies ("No compression", "Expected Attention (50%)", and "Expected Attention (90%)") as sequence length increases from 10,000 to 120,000. Memory usage rises monotonically for all strategies, with "No compression" consistently consuming the most memory and "Expected Attention (90%)" the least.
### Components/Axes
- **X-axis**: Sequence Length (10,000 to 120,000 in 20,000 increments)
- **Y-axis**: Peak Memory Usage (GB) from 10 to 45
- **Legend**:
- Blue: No compression
- Red: Expected Attention (50%)
- Brown: Expected Attention (90%)
- **Bar Groups**: Three bars per sequence length, grouped by compression strategy
### Detailed Analysis
1. **No compression** (blue):
- Starts at ~17 GB (10,000 seq) and rises to ~43 GB (120,000 seq)
- Linear increase with ~0.18 GB per 1,000 seq increment
2. **Expected Attention (50%)** (red):
- Starts at ~16 GB (10,000 seq) and rises to ~35 GB (120,000 seq)
- Linear increase with ~0.16 GB per 1,000 seq increment
3. **Expected Attention (90%)** (brown):
- Starts at ~15 GB (10,000 seq) and rises to ~30 GB (120,000 seq)
- Linear increase with ~0.13 GB per 1,000 seq increment
### Key Observations
- Memory usage scales linearly with sequence length for all strategies
- Compression reduces memory usage by ~6-7 GB per 10,000 seq increment
- The memory savings between 50% and 90% compression diminish at higher sequence lengths
- No compression uses 2.5-3x more memory than 90% compression at maximum sequence length
### Interpretation
The data demonstrates that memory requirements grow proportionally with sequence length, with compression strategies offering predictable reductions. The diminishing returns of higher compression (50% vs 90%) suggest that beyond a certain point, additional compression effort yields smaller memory savings. This implies a potential cost-benefit tradeoff for systems handling large sequences: while 90% compression saves memory, the computational overhead of achieving higher compression ratios may not justify the marginal gains. The consistent linear trends indicate no unexpected memory spikes or anomalies in the tested range.