The cheaper model made things up

NTULearn toolkit

Choosing a vision model by testing both on the same 99 slides.

NTULearn toolkit pipeline

Frames sampled from a lecture recording are fingerprinted, and near-duplicates are dropped, leaving about one image per slide. A vision model reads those into notes, while the transcript feeds an evening digest.

The model was picked by running both candidates over the same 99 frames and checking every disagreement against the slides themselves.

The NTULearn toolkit is a tool I built for my own studies. It downloads lecture recordings and turns them into study notes: frames are sampled every 5 seconds, near-duplicates are dropped using perceptual hashing, and a 2-hour lecture becomes about 100 slide images that a vision model reads into notes.

problem

Once a lecture is processed, the tool deletes or shrinks the original video to save space, so the notes become the permanent record. If the model hallucinates a slide, there is no video left to check against.

A cheaper model is tempting for a job that runs every night, but I needed to know whether it was accurate enough.

solution

I ran the same 99 frames from a real 2-hour lecture through both Sonnet and Opus and compared the outputs. They disagreed on 7 frames about what was on screen.

Checked against the actual images, Opus was right: Sonnet had invented a File Explorer listing on one frame and mistaken PowerPoint for RStudio on another. Opus also captured about 19% more content.

I chose Opus, because here accuracy matters more than cost.