Home / Articles / Blog

How to Understand LLM Output Faster: Four Formats, from Controlled Text to Explainer Videos
LLMAIManimRemotionShowtimeDeveloper tools

How to Understand LLM Output Faster: Four Formats, from Controlled Text to Explainer Videos

Andrej Karpathy's four output formats for understanding LLM work—ASD-STE100, diagrams, interactive HTML, and explainer videos—with prompts, Manim and Remotion paths, and Showtime for local agent-rendered video.

Shrey Vijayvargiya

AI now writes the code, drafts the report, and runs the research. That moves your job. You do less of the doing and more of the checking, and checking only works if you understand what the model gave you.

Most people read the answer as plain chat text and move on. Andrej Karpathy posted on X that text is often the weakest format for this job. He lists four output formats, each one easier to absorb than the last: controlled-language writing, diagrams, web pages, and custom explainer videos. This guide explains each format, shows the prompts to use, walks through the video workflow step by step, and points at Showtime—a local video studio your coding agent can run with templates and examples you can copy.

Why format matters now

As models do more of the work, humans move up into oversight and understanding. The model handles the legwork. You decide whether the result is right. That decision needs comprehension, not just a green checkmark.

The second claim is about cost. Code is now cheap, so you can ask for large, one-off software that was never worth building before: a throwaway web app that explains one bug, or a video that explains one concept. You use it once and delete it.

Put together, the advice is simple. Do not accept the default format. Ask for the format that your brain reads fastest, and keep pushing toward richer formats when the topic is hard. If you are already thinking in terms of wrappers around models, our agent harnessing guide is the same idea one level up—the model is not the whole product, the interface to it is.

Start with plain, controlled English (ASD-STE100)

ASD-STE100 is a controlled version of English that was built so that technical instructions cannot be misread. The standard comes from the aerospace industry. Issue 9, published in January 2025, has 53 writing rules in nine sections that cover word choice, grammar, sentence structure and style, plus a dictionary of approved words.

The core idea is the same one that Skybrary describes: one word has one meaning, and where English has several synonyms, the standard picks one. The original goal was to help readers who speak English as a second language follow maintenance manuals without guessing.

This helps with LLM output for a plain reason. Models tend to write long sentences full of hedges. A strict rulebook is a constraint they follow well. One write-up on the topic notes that people began packaging STE as an agent skill in 2026 to force models into this style.

Prompts that actually work

Name the standard in your prompt:

Explain how OAuth 2.0 works in ASD-STE100. Use short sentences, active voice, and one instruction per sentence.

The full specification is strict. If the result feels stiff, ask for a softer version:

Explain database indexing about 80% of the way to ASD-STE100. Keep every sentence under 20 words.

Two limits to know. The official standard says full compliance needs its own dictionary, so a model gives you the style, not a certified document. And STE was designed for procedures and descriptions, so it works best for "how does this work" and "what do I do" questions, not for opinion or analysis.

When a picture beats a paragraph

A diagram shows structure that paragraphs hide. Request flows, state changes, dependencies and architecture are all spatial. Prose forces you to build the picture in your head. A diagram gives you the picture and lets you check it.

Prompts that work:

Show the request lifecycle of a Next.js app as a flow diagram.

Draw the auth flow between browser, API and database as a sequence diagram.

If your tool cannot render images inline, ask for Mermaid or SVG code. Mermaid is plain text that many editors, wikis and documentation sites render directly. SVG opens in any browser.

A diagram also makes errors visible. A missing arrow or a loop that goes nowhere is easier to spot than a missing sentence. That is the oversight benefit: you can audit the model's understanding of the system at a glance.

One HTML file you can click through

HTML turns an explanation into something you can click. Models are now good at front-end work, so a single self-contained page can include animation, step-through controls, sliders and toggles.

The key word is interactive. A text explanation of JWT authentication is static. A page where you click through each step, change the token expiry and watch what breaks teaches the idea faster.

Prompts that work:

Explain how JWT auth works as a single-file HTML page. Add a step-by-step walkthrough, an animated token flow and a toggle for "token expired".

Build an HTML page that explains Big-O with a slider for input size and a live chart of the growth curves.

Open the file in a browser, or publish it as a link if your tool supports that. Then keep asking for changes: "add a failure case", "show the data at each step", "make the dark theme the default". Each round costs seconds.

If you build for the web, you can reuse the same page for a blog post or a product walkthrough. Browse the iHateReading blog index for step-by-step front-end topics that pair well with this approach, or ship a directory listing with the same disposable mindset in build a directory website that earns online.

The heavy lift: explainer videos

A custom explainer video is the richest output format, and it is starting to work. Karpathy calls it the format he is most bullish on: a bespoke video on any topic, in the style of 3Blue1Brown, with narration from a voice API.

A video is four jobs. Each one has a clear tool.

JobWhat does itTool options
Script and scene planThe LLMAny capable model
AnimationCode that renders framesManim (Python) or Remotion (React)
NarrationText-to-speechElevenLabs API, or local options such as Kokoro
AssemblyMerge audio and videomanim-voiceover or ffmpeg

We have walked through programmatic video before in Prompt to Video App—Remotion on the server, plain English in, MP4 out. The formats below are the same family of idea with more control per scene.

Manim or Remotion?

Manim was created by Grant Sanderson of 3Blue1Brown. It uses Python and suits mathematical and scientific animation. Remotion uses TypeScript and React, and each frame of the video is a React component. The same comparison names Remotion as the better fit for web developers who want scalable video pipelines.

A simple rule: choose Manim for equations, graphs and "how does this algorithm work" topics. Choose Remotion if you already write React and want product demos, dashboards or data-driven videos.

You do not have to wire the tools together alone. A community toolkit for Claude Code collects skills and plugins that cover both Manim and Remotion, including a plan-write-render-iterate loop.

Voice: hosted or local

You have two paths. ElevenLabs is a hosted service with polished voices and no local setup. Kokoro is an open model with downloadable weights under the Apache-2.0 license, so you can run and inspect it yourself. Pick the hosted service for speed and voice quality. Pick the local model for cost control and privacy.

Manim Voiceover is a plugin that adds voiceover directly in Python, so you do not need a video editor. It can also time animations to specific words in the narration, using Whisper. It supports several speech services, including free ones such as gTTS and pyttsx3.

Eight steps, low quality first

  1. Install the tools. You need Python, ffmpeg and Manim Community Edition.
    bash
    pip install manim
    pip install "manim-voiceover[gtts]"
    Use the manim package name. A separate version of Manim exists for 3Blue1Brown's own videos, and it installs under a different name.
  2. Ask the model for a scene plan. Request 5 to 8 short scenes. For each one, ask for the narration text and a description of what appears on screen.
  3. Generate one scene file at a time. Do not ask for the whole video in one prompt. Small scenes fail in small ways.
  4. Render at low quality first.
    bash
    manim -pql scene.py AttentionIntro
  5. Paste errors back to the model. Give it the full traceback and let it fix the code.
  6. Add narration. With Manim Voiceover, a scene looks like this:
    python
    from manim import *
    from manim_voiceover import VoiceoverScene
    from manim_voiceover.services.gtts import GTTSService
    
    class AttentionIntro(VoiceoverScene):
        def construct(self):
            self.set_speech_service(GTTSService())
            title = Text("Attention", font_size=56)
            with self.voiceover(text="Attention lets a model look at every word at once.") as tracker:
                self.play(Write(title), run_time=tracker.duration)
  7. Merge separate audio if needed. If you made narration with an outside API such as ElevenLabs, join it to the video with ffmpeg:
    bash
    ffmpeg -i scene.mp4 -i narration.mp3 -c:v copy -c:a aac -shortest final.mp4
  8. Render at high quality once the low-quality version looks right.

A starter prompt:

Create a 3Blue1Brown-style explainer on how attention works in transformers. Write it as Manim scenes. Use at most 6 scenes of 20 to 40 seconds each. Give a narration line for each scene and use manim-voiceover. List the setup commands and the file structure first.

If you prefer React, the Remotion route is similar: install Remotion, describe each scene as a component, preview in Remotion Studio and render to MP4 from the command line.

Showtime: your agent directs, your machine renders

Manual Manim and Remotion workflows teach you what is happening under the hood. When you want the whole pipeline in one sentence—script, voice, music, captions, QA, and exports—use Showtime.

Showtime is a local video studio for your coding agent. You describe a video in plain language; the agent plans scenes, renders on your machine, and drops a folder with final.mp4, poster, captions, and platform exports. No cloud AI API keys and no uploads of your source material to a vendor. It is built for Claude Code, Codex, Cursor, Devin, OpenCode, and any agent that supports Agent Skills.

Quick start in Claude Code:

text
/plugin marketplace add FavioVazquez/showtime
/plugin install showtime@showtime

Then ask:

text
Make a 20-second launch video for this repo, with a voice-over and upbeat music.

The first run downloads models and tools into ~/.showtime (on the order of 0.6–0.8 GB, usually a few minutes). After that, each job lands in showtime-out/<name>-<timestamp>/ with final.mp4, poster.jpg, share.txt, and exports/. Say "show me options first" to open a local studio board before anything renders; say "no crew" to keep the whole job in one session.

Templates and examples you can steal. The companion repo showtime-examples ships 22 finished projects with prompts, project files, and full-quality videos: launch clips, heat-pump explainers (English and Spanish), tutorials recorded from a real app, vertical shorts, beat-synced montages, repo release videos, Manim math explainers, and more. Copy a prompt from the table in that repo— for example the 70-second Manim circle-area explainer—and paste it into your agent.

Showtime also closes the loop with Format 3. Motion scenes are HTML, CSS, and canvas rendered frame-exact in headless Chrome; math uses Manim; you can export a single offline HTML video with chapters and keyboard shortcuts:

bash
showtime export html

Several gallery examples include those HTML players alongside MP4—useful when you want shareable motion without asking someone to download a file.

For repo and paper explainers, release videos from changelogs, and brand-aware launch films, the Showtime docs and showtime guide commands are the source of truth. Treat Karpathy's four formats as the what; Showtime is one how when the format you need is video and you would rather direct than wire ffmpeg yourself.

Pick the format before you pick the tool

Do not jump straight to video. Each step up costs more time, so match the format to the difficulty of the topic.

  • Simple fact or procedure: controlled-language text.
  • System with parts and connections: a diagram.
  • Concept that depends on "what if I change this": an interactive HTML page.
  • Hard idea with motion, math or a sequence over time: an explainer video—by hand with Manim/Remotion, or with Showtime when you want the agent to run the studio.

If one format does not click, regenerate in another. The output is cheap, so treat it as disposable.

Throwaway explainers are the point

The second half of the post carries the bigger idea. When code is abundant, a custom tool for one question becomes reasonable. A throwaway page that visualises one log file is now a five-minute job. Before, nobody would have built it.

For indie developers and small teams, this changes where time goes. Instead of reading a dense pull request, ask for a visual walkthrough of the change. Instead of reading a long spec, ask for an interactive version with the edge cases highlighted. You spend your effort on understanding, which is the part that stays with you. If you are turning that traffic into users, pair explainers with directory listings—see how to get users for free by submitting your SaaS.

Pretty output still lies

A good-looking format raises one risk. A smooth video with a confident voice can still explain something wrong, and polish makes errors easier to trust.

Use these checks:

  • Read the narration script before rendering. Errors are cheap to fix in text.
  • Verify numbers and claims against the source. Do not trust a chart because it looks clean.
  • Test interactive pages with edge cases. Set the slider to zero, set it to the maximum, and see if the page still makes sense.
  • Ask the model to list its assumptions. Then check the list.

The format helps you understand. It does not replace judgment. For keeping prose tight without sounding like generic AI filler, making AI applications without slop is a useful companion read.

For more practical AI and developer guides in this style, read the monthly magazine or browse the full article index.

External Links (14)

Explore topics

More in

Weekly letter

Subscribe to iHateReading

Our once-a-week newsletter on programming, jobs, AI, and building products online.

Our once a week newsletter on Programming, Jobs, AI, and Business