Vision model against OCR

Claim. On strict verbatim phrases, a local vision model reproduced 31 of 32, Tesseract 22, RapidOCR 16.

Method. benchmark_vlm.py on real frames from the target: 3 BIOS screens and the console line, 32 phrases. A phrase counts only if reproduced verbatim (whitespace normalised, case kept). The model was qwen3-vl 30B-A3B through Ollama at temperature 0. Tesseract 5.5.3, RapidOCR 3.9.2.

Sample size. 4 frames, 32 phrases, 1 vision model. This is small.

Result.

EngineVerbatim phrasesSeconds per frame
Tesseract 5.5.322/32about 1
RapidOCR 3.9.216/32about 1.3
qwen3-vl 30B-A3B31/326 to 26

Repeat runs of the model gave identical text (3 runs on one frame, 2 on each of four). It read the slashed zeros correctly, which both OCR engines get wrong. Its one miss was turning @ into 0. It returns no boxes, so it cannot place a click.

Limits.

Design consequence. OCR for speed, boxes and waits; a vision model as an exact reader for a screen that matters; a phrase list or regular expression decides on disagreement.

Raw data. docs/evidence-sources/20260929-llm-kvm-landscape.md (section "What we measured") and scripts/kvm/benchmark_vlm.py in this repository. Frames are not published.

Original record: read the Markdown source.