Back to all posts
Training·9 min read

Yes, Claude Can Make a Training Guide. Now What?

Claude built a genuinely excellent video guide. Lattify did it faster, cut the video more reliably, and handed me something I could assign that afternoon.

E

Eamonn Best

Founder, Lattify · July 24, 2026

Yes, Claude Can Make a Training Guide. Now What?

Every time I explain what I'm building, I get the same question about ninety seconds in, delivered the way you'd tell someone their shoelace is undone.

"Can't ChatGPT just do that now?"

It's a fair question, and this year the answer stopped being an easy no. So I ran it properly: the same store opening-and-closing video through ChatGPT and Claude, and through Lattify, and then I sat with what came back.

How I scored it

The video was eight minutes and forty-four seconds of an ordinary shop open and close, filmed on a phone, which I had permission to use. Nothing staged, nothing tidied.

I built a 38-point reference audit covering the procedural instructions and operational details I wanted to test, from lifting the floor bolt to pushing the door latch down so you can't lock yourself out. It is a completeness check, not a prescription for 38 buttons. Every system was allowed to combine related instructions into sensible groups, three light switches under one "turn the lights on," four bag types under "the bags," the way any decent trainer would.

A detail counted when it was stated directly or sat inside an accurately labelled group. Raw footage playing under an unrelated or fabricated label did not count. Timing was scored at the group level: did the control land near the beginning of everything its group promised, and did it stop after the last relevant instruction? Grouped instructions weren't penalised for lacking separate buttons.

Two honest caveats. One video, one kind of shop. And the timing pass wasn't blinded, because I'd already seen the guides by the time I went back through the audio and footage to mark fair start and end windows.

Round one: what an ordinary owner would ask for

The first prompt was the one a shop owner would actually type. Make a complete, usable video guide from this video, with clear steps and clickable timestamps that jump to the relevant moment, suitable for a worker learning the job. One upload, one go.

ChatGPT came back with a handsome interactive page, and about seven of the 38 reference details survived in usable form. It had skimmed the video and written a confident, generic retail guide out of what it already knew about shops, adding a pre-entry damage inspection that never happens. Run again, it got worse and invented an alarm procedure.

Claude Fable 5 did far better. Its first run had genuinely understood most of the shop. The trouble was navigation: it gave start-only controls, several of them landing a long way from the thing they promised, and nothing stopped, so the video just rolled on into the next job. Then I ran it a second time, and it handed me a guide to a café. An espresso machine, a food display, an A-board for the pavement. There is no café in that video. It's a clothing shop, and what I was reading was a complete, polished, plausible procedure for a business that does not exist, delivered with the same confident tone as the good run.

Grouping related actions together is fine, and Lattify does it too. What let the Fable guide down was inaccurate navigation, controls with no stopping point, and a second run that came back completely different.

The ordinary-prompt test showed the problem, but not the ceiling. Claude had understood most of the shop. I wanted to know what a frontier model could do if I told it exactly what a proper instructional player needs.

Round two: the proper brief

So I made the fight harder for Lattify. I wrote the prompt a person who knew what they wanted would write: complete coverage, sensible navigable groups, an accurate start with minimal lead-in, an accurate end after the last instruction, and playback that pauses on its own at the end of each step. Nothing in it demanded that every reference detail become its own button. Sensible grouping was allowed.

I switched from Claude Fable 5 to Claude Opus 4.8, in a fresh conversation. This was not an upgrade in raw model capability, since Anthropic positions Fable 5 above Opus in general capability, but Opus produced the stronger artifact after the more explicit brief and an additional turn.

This time Claude was excellent. Not "excellent for an AI experiment." Excellent enough that a shop could genuinely use it. It covered the whole procedure: an attractive guide with thirty-eight controls, real start and end times, automatic pausing, the tools and materials and warnings pulled out, a working copy of the video packaged in. As a one-off guide, more than good enough.

ChatGPT got the same explicit boundary and auto-pause requirements in a fresh conversation. It delivered eighteen controls in 5:58. Nine of them could be mapped to the real procedure, covering eleven reference details. Only two landed usefully, and none correctly bounded the material their labels promised.

The scorecard

Two things get counted below. Content covered is whether an instruction made it in at all, inside an accurately labelled group. The landing and boundary rows are scored only over each system's source-mappable instruction groups, the controls that correspond to a real part of the procedure.

Explicit-boundary, group-aware resultLattify repeatClaude Opus 4.8GPT-5.6 Sol Medium
Content covered38 / 3838 / 3811 / 38
Delivered controls303818
Source-mappable groups / delivered controls30 / 3033 / 389 / 18
Useful landing29 / 30 (97%)30 / 33 (91%)2 / 9 (22%)
Correct whole-group boundary26 / 30 (87%)21 / 33 (64%)0 / 9 (0%)
Auto-pauseDefaultExplicitly requestedExplicitly requested
Native assignment and completion trackingYesNoNo

Timing denominators use source-mappable instruction groups, not every delivered control. All thirty Lattify controls mapped to the source. Claude's thirty-eight controls became thirty-three comparable groups after adjacent controls that split one instruction were combined and two non-comparable controls were excluded. Nine of ChatGPT's eighteen controls mapped to the source, covering eleven reference details. Unmappable controls were excluded from timing rather than automatically counted as timing failures, so this treatment favours the generalist models rather than Lattify.

A word on the clocks, because they aren't all the same. Lattify reached a ready guide in three minutes and sixteen seconds after I pressed Create, with roughly forty seconds of upload before that, so about three minutes fifty-six from picking the file to a finished guide. It ran twice and hit that mark both times; the table uses the untouched repeat run. ChatGPT took five minutes and eight seconds after the prompt, closer to five fifty-eight counting its upload. Claude Opus was a two-turn workflow spread across a conversation.

Lattify and Opus both captured the complete procedure. There is no honest basis for saying Lattify understood more of this video. The difference was consistency and what surrounded the guide: Lattify used fewer groups, landed and bounded them more reliably, and finished in minutes.

Lattify was not flawless. The raw guide was complete and structurally usable, but I would remove the unsupported doorstopper wording in its closing control before publishing it. The important difference is that Lattify gave me an editor and a managed update path for doing exactly that.

What happens on Monday

We do not need to imagine the next model winning. Opus already produced a guide that was more than good enough for this shop.

Congratulations. You have an HTML file.

And that is where the real product question begins.

The HTML file can't assign itself to the two new starters coming in on Monday, or tell you whether they finished it. It can't become tonight's closing checklist, ask for a photo as proof, translate itself for the closer who reads Spanish, or add subtitles and a slower playback speed for someone still learning. It has no hands-free mode for a chef with both hands in a bowl, no employee app or WhatsApp line where a worker can ask a question mid-shift and get a grounded answer from your own training, no search across the rest of your guides, no analytics on where people got stuck, and no managed way to update it when the process changes in September.

You could build all of that. Opus just showed you can build a surprising amount. But now you are building and running your own training platform on top of running the shop. What is your time worth again?

Lattify's guide was already inside that system. I could preview it, fix the doorstopper wording, assign it to staff, or turn it into an opening and closing checklist that same afternoon. That is what execution-ready means.

The models will keep getting better

Frontier models made Lattify possible, and their getting better is good news for me. Lattify is built out of them: speech recognition, visual analysis, and specialised language models, each doing the part it is strongest at. As those parts improve, so does what Lattify puts out.

This recording barely tested Lattify's visual-analysis advantage, because the narrator explained nearly everything aloud. A proper multimodal benchmark would include tools, controls, labels and demonstrated actions that are visible but never spoken. That test has not been run yet, so I am not claiming it.

Claude can build an excellent guide. Lattify can produce an equally complete one in minutes, with more reliable navigation, already living inside the system that gets the work assigned, completed, checked, searched, translated and maintained.

The guide is becoming a commodity. Execution is not.

If any of this sounded familiar, we built Lattify for exactly this problem.

Join the Waitlist