# Sources and figure data

## Kernel performance

The FP8 comparison uses matched GB200 measurements of Dense (cuDNN), Sol-Attn1,
and Sol-Attn2. Each measurement covers the complete QKV-to-output call,
including input-dependent preparation. These are operator timings, not model
end-to-end inference speedups. The benchmark uses 32 heads, head dimension 128,
32K/64K/128K tokens, and 80/85/90% sparsity.

The measured FP8 kernel has 64 compressed slots. The Stage A1 model uses the
Q32 complementary-tail approximation described in the method; the two figures
report different implementations.

## Stage A1 curves

VSA, SLA, and Sol-Attn2 each completed 200 updates at fixed 90% token-pair
sparsity with matching training inputs. Validation uses four held-out cases,
eight noise intervals, and 50 layers, at steps 0/25/50/100/150/200.
The plotted Sol-Attn2 is the accepted Q32 variant with the lowest final mean
paired normalized loss among the compared Sol variants. Trainable parameter
counts are shown with each method. The loss is layer-output error relative to
the dense teacher, not a generated-video quality metric.

- [Figure values and configuration](data/blog-data.json)
- [Complete plotted curves](data/curves.json)
- [Editable manuscript](draft.md)

## Related work and attribution

The ExpCast discussion cites Xingyang Li et al., *VC-Attention: Value Smoothing
and Softmax Casting for Low-bit Attention*,
[arXiv:2609.15810v1, Section 3.3](https://arxiv.org/html/2609.15810v1), CC BY 4.0.
The displayed expression is the unit-scale form used by our implementation;
it approximates exponential probabilities.

The research-page presentation was informed by the
[VDN page](https://openvdn.github.io/) and the
[Sol-H3-Spark page](https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3-Spark/).
Their branding and results are not reused as ours.

## Draft status

The abstract's end-to-end speedups and post-QAD quality statement are pending
result fields. The four Dense/Sol-Attn2 sample cards are placeholders: one before Method and three at the end.

## Output-composition notation

The table describes one attention head before head concatenation and the O output layer. Each method selects its own E_i and normalizes o_s over those tokens. SLA learns an affine projection (with bias) of all-token linear attention using channel-wise softmax features. VSA combines sparse attention with all-block pooled attention through an unconstrained, bias-free linear gate on the input hidden state, with one coefficient per output channel. Its pooled scores also rank blocks for routing. Sol-Attn2 instead mixes selected attention with the complementary tail according to their masses. All three A1 runs also train QKVO LoRA.
