node_58b5d013 · 2026-09-27 21:19 UTC · Chrome on Android
First Adreno 750 on the network
The reference machine scores 1,000. This one delivers 0.38× its throughput.
Streaming loads. The path every token takes through the weights.
Streaming stores. KV-cache appends and activations.
Buffer to buffer. Read and write traffic together.
Scattered gathers. Embedding lookups and sparse access.
Dense matrix multiply. The arithmetic inside every layer.
Continuous load. What the machine holds once it heats up.
No fixed published peak exists for this class, so it is ranked against its own runs only.
- Llama 3.2 3B3.2 GB≤ 29.0 tok/s
- Llama 3.1 8B6.4 GB≤ 11.6 tok/s
- Llama 3.3 70B45.9 GB≤ 1.3 tok/s
- Phi-410.7 GB≤ 6.3 tok/s
- Gemma 2 27B18.7 GB≤ 3.4 tok/s
- Qwen 2.5 32B22.2 GB≤ 2.8 tok/s
4-bit weights, 8K context. The ceiling is read throughput divided by the weights read per token; real runtimes land below it. Memory size is unknown for this device, so fit is not judged.
- Verified run+100
- First run on Adreno 750+100
- Early data for this class+250
- Genesis wafer+500
Deep report
Open during GenesisEvery phase sample by sample, the spread between the 5th and 95th percentile, this run against its class median, and a full export.
How does your machine compare?
Forty-five seconds, in this browser.