0:00 / 0:00
Ollama on an M1 Pro — llama3.2:1b
llama3.2:1b · 86.6 tok/s decode · 1 stream
A Mac card that now names the machine: toktape reads the host line from sysctl where there is no /proc.
The GPU slot stays empty on purpose. Apple Silicon has no VRAM of its own, and a VRAM bar over a unified pool would be a claim about memory that does not exist.
decode
86.6tok/s
prefill
?tok/s
ttft p50
4059ms
| model | llama3.2:1b |
| quantisation | ? |
| size | ? |
| engine | openai |
| host | macos |
| streams | 1 |
| workload | 6147 in / 1048 out · cache 0% |
| context | ? |
| config | ? |
| prompt set | prompts@v2 |
| recorded | 2026-09-21T01:32:28Z |
| toktape | dev |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 5 caveats —
client_timed, recorded, recorded, recorded, no_proc_view. The card that comes with this record spells each one out.Page views are counted with Cloudflare Web Analytics: no cookies, no fingerprinting. The toktape binary itself sends nothing anywhere.
