Home / Articles / Blog

Tubelog: Turn Any YouTube Channel Into a Markdown Library and AI Blogs (Free, Open Source)
YouTubeTubelogMarkdownAIOpen source

Tubelog: Turn Any YouTube Channel Into a Markdown Library and AI Blogs (Free, Open Source)

Tubelog is a free, open-source local tool that archives a YouTube channel as Markdown files and can write an AI blog from each transcript.

Shrey Vijayvargiya

Every few weeks, someone like Andrej Karpathy or a Y Combinator channel posts an hour of dense, practical material. Most of it never becomes text. It stays locked inside a video that you have to scrub through to find one idea.

That is a content gap, and it is also a content engine. YouTube is the largest pile of expert explanation on the internet, and the raw material is sitting right there as captions. I built Tubelog, a free, open-source "YouTube CRM," to turn that pile into a local Markdown library you can search, study, and write from.

This post covers what Tubelog does, how it works, how to run it in about two minutes, and how to turn a transcript into a blog post that is actually worth publishing. There is a copy-paste checklist near the end.

What is Tubelog?

Tubelog is a local tool that archives YouTube channels and videos as Markdown files, one file per video, and can optionally generate an AI-written blog from each transcript. It runs as a CLI, an HTTP API, and a small dashboard on your own machine.

The mental model from the repo is short: the video is the source, the transcript is the raw knowledge, the blog is the lesson, and the filesystem is the database.

What it does, in plain terms:

  • Fetches videos from a channel. Give it a channel URL, an @handle, or a channel ID.
  • Fetches the transcript for a video. Give it a single video URL or ID.
  • Stores each one as Markdown in a videos/<channel>/<video-id>.md folder structure.
  • Generates a blog on request. An LLM routed through OpenRouter writes the post from the transcript, only when you ask for it.

The code is on GitHub under the MIT license.

Why YouTube is the best content engine for blogs

YouTube works as a content engine because experts explain things there at length, in their own words, and for free. Karpathy's Zero to Hero course walks from backpropagation up to GPT-style models, step by step, in code. His 2025 "Deep Dive into LLMs like ChatGPT" runs over three and a half hours. That is a book's worth of explanation, and most people will never sit through all of it.

A blog fixes that for the reader. It is skimmable, searchable, and linkable. And for you as a builder, it works the other way too: if you keep the transcripts of the channels you learn from, you have a private, greppable knowledge base instead of 40 browser tabs.

The same logic powers the research layer behind this site. In how I built an agent system that researches 100 blogs a day, YouTube transcripts are one of four input streams. Tubelog is that one stream, extracted and turned into a standalone tool.

How Tubelog works

Tubelog has one core and three front doors. The logic for talking to YouTube, fetching transcripts, saving files, syncing, and calling the LLM lives in a handful of modules under src/. The CLI, the Hono API, and the React dashboard all call those same modules, so they never drift apart.

text
YouTube channel or video
        ↓
metadata + transcript
        ↓
videos/<channel>/<video-id>.md
        ↓
optional OpenRouter blog
        ↓
dashboard, CLI, and API

The tech stack:

  • Frontend: Vite, React, Tailwind, and shadcn/ui
  • Backend: Hono, a lightweight framework built on Web Standards that runs on Node.js, Bun, Deno, and Cloudflare Workers with the same code
  • AI routing: OpenRouter, a single API gateway for hundreds of models where switching models is a one-line change

Quick start: run Tubelog in two minutes

  1. Clone the repo: git clone https://github.com/shreyvijayvargiya/tubelog
  2. Install and configure:
bash
npm install
cp .env.example .env
npm run dev
  1. Open the dashboard at http://localhost:3000.

To archive a whole channel from the command line without spending any AI credits:

bash
npm run tubelog -- sync https://www.youtube.com/@fireship

To archive it and write blogs for the new transcripts:

bash
npm run tubelog -- sync https://www.youtube.com/@fireship --ai

For a single video:

bash
npm run tubelog -- transcript https://www.youtube.com/watch?v=VIDEO_ID
npm run tubelog -- blog VIDEO_ID

"No APIs, no auth, no databases": what that actually means

Tubelog needs no YouTube API key, no user accounts, and no database. That is the "free" part, and it is worth being precise about, because one piece still needs a key.

  • No YouTube Data API. Tubelog reads the public channel page and pulls captions the way a browser would, with a few fallback methods.
  • No auth. It is a local tool. There are no logins or sessions.
  • No database. Markdown files on disk are the storage layer. Sync is idempotent because the video ID is the filename, so running it twice never creates duplicates.
  • One optional key. The AI blog step needs an OPENROUTER_API_KEY in your .env. Archiving transcripts works without it, and the server never calls OpenRouter unless you pass --ai or set generateBlog: true.

A filesystem as the database is also a good default for a solo project in general. It costs nothing to run, it diffs cleanly in git, and any tool that reads text can use it. The same instinct, keep infrastructure boring and cheap, is behind how we keep ihatereading.in under $100 a month at 10k visitors.

What a Tubelog Markdown file looks like

Each video becomes one file with frontmatter and the transcript:

markdown
---
video_id: "No-JPdFvYWU"
channel: "Fireship"
title: "How AI Agents Work"
url: "https://www.youtube.com/watch?v=No-JPdFvYWU"
published_at: "2026-10-01"
transcript_available: true
blog_generated: false
---

# Transcript

Full transcript here.

When you generate a blog, it is appended to the same file under an "AI Generated Blog" heading, followed by a small metadata block recording the model and timestamp. One video, one file, and the original watch URL stays in the frontmatter so you can always link back to the source.

Plain Markdown means the library works with everything: your editor, grep, a static site generator, or an LLM that you point at the folder.

AI blog generation with OpenRouter

Tubelog generates blogs through OpenRouter so you can swap models without touching code. The default is google/gemini-2.5-flash, and you can override it with the OPENROUTER_MODEL environment variable or in tubelog.config.js. OpenRouter is OpenAI-SDK compatible, so pointing at a different model really is a string change.

Tubelog supports a few writing styles: educational, tutorial, explainer, technical, and beginner. You can also toggle takeaways, an embedded link to the original video, and code examples. The system prompt lives in a plain file, prompts/blog-writer.md, which tells the model to teach the ideas, stay inside the source, and keep only code that the transcript actually contains. Edit that file and you change the voice of every blog.

A prompt, a model, and a loop around it is a tiny version of what an agent harness does at larger scale. The model matters, but the wrapper around it matters just as much.

From transcript to a blog worth publishing

A transcript is raw material, not a draft. Several sources on YouTube repurposing make the same point: publishing a transcript with light edits is the most common mistake, and the better approach is to find the three to five main ideas, make those your headings, and write each section properly using the transcript for substance rather than phrasing.

Spoken language is also messy on the page. Pasting captions straight into a post gives you broken phrases, repeated clauses, and fragments. One guide on turning transcripts into posts estimates that a roughly 3,000-word transcript condenses into a 1,000 to 1,500-word post that reads better.

Here is a workflow that treats the AI output as a first draft:

  1. Pick videos worth the effort. Not every video deserves a post. Prefer tutorials, teardowns, and talks with a clear takeaway.
  2. Generate the first draft with Tubelog. Use the tutorial or technical style for developer content.
  3. Rewrite the structure. Cut intro chatter and filler. Reorder sections so each one makes a claim in its first sentence.
  4. Add what the video could not carry. Code you tested yourself, a comparison table, a diagram, a mistake you hit, a link to the docs.
  5. Credit and link the source. Name the creator, link the video, and embed it. Your post should send people to the original, not replace it.
  6. Fact-check anything specific. Names, numbers, versions, and commands. LLMs and auto-captions both get these wrong.

The scale trap

Because Tubelog can archive a whole channel in one command, it is tempting to generate a hundred posts and publish them all. Don't. Google's scaled content abuse policy targets many pages created mainly to manipulate rankings rather than help readers, and it applies whether the pages come from AI, humans, or a mix.

The way to stay on the right side of that is editorial: fewer posts, each with real added value and a human review step. That is the same principle behind how we publish 100+ SEO blogs a day for under $10 without tripping the scaled content penalty. Automation handles the research and the draft, and a person handles the judgment.

Limits and gotchas

Being honest about what Tubelog does not do:

  • Missing captions. If a video has no captions Tubelog can read, the file is still saved with transcript_available: false and the sync keeps going. Run sync --force to retry later.
  • Rate limiting. YouTube sometimes blocks automated requests with a 429 or a consent wall. Tubelog treats this as retryable rather than permanently missing, and it adds a short delay between requests.
  • 20 videos per sync by default. That is the sync.maxVideos setting. Run sync again to continue through the channel, since already archived videos are skipped.
  • Shorts are skipped.
  • It is a local tool. There are no accounts, billing, cloud storage, or scheduled sync in this version. MCP and Claude integration, cloud storage, and scheduled sync are on the roadmap.
  • Respect creators. A transcript is someone else's work. Use it to learn, to summarize, and to write a new piece with your own analysis, and always credit and link the source. This is general guidance, not legal advice, so check the creator's terms if you plan to republish at scale.

What to build on top

Tubelog is deliberately small, so it is a good base for your own ideas:

  • A searchable personal wiki. Sync the 10 channels you learn from and use tubelog search or point an LLM at the videos/ folder.
  • A content-ideas pipeline. Sync a channel, scan titles and transcripts for angles, and queue the best ones for human-written posts.
  • A creator tool. If you make videos, archive your own channel and turn each upload into a companion blog post with the video embedded.
  • An agent knowledge source. The Markdown library is a ready-made context folder. If you are learning how agents use memory and tools, the AI Agent Developer Roadmap covers frameworks, MCP, memory, and deployment.
  • A hosted SaaS. The dashboard is just a client of the API, so you can host the Hono server and add accounts. If you ship it, list it on the SaaS directories we maintain.

Free checklist: YouTube to blog in 10 steps

Copy this into your notes:

  1. Choose one channel you trust and one playlist or video with a clear takeaway.
  2. Run tubelog sync <channel> without --ai first, and read the transcript.
  3. Write the one-sentence claim the post will make.
  4. Pull 3 to 5 headings from the main ideas, not from the video's chronology.
  5. Generate a draft with tubelog blog <video-id> in tutorial or technical style.
  6. Delete filler. Make each section open with a clear statement.
  7. Add something the video lacked: tested code, a table, a diagram, or a correction.
  8. Verify every name, number, version, and command.
  9. Credit the creator, link the video, and embed it.
  10. Add 2 to 3 internal links, publish, and move on to the next video.

FAQ

Is Tubelog free? Yes. It is open source under the MIT license and runs locally. Archiving transcripts costs nothing. Only the optional AI blog step uses credits on your OpenRouter account.

Do I need a YouTube API key? No. Tubelog does not use the YouTube Data API. It needs an OpenRouter key only if you want AI-generated blogs.

Can Tubelog download all transcripts from a YouTube channel? Yes, in batches. It archives up to 20 new videos per run by default, skips ones it already has, and lets you run sync again to continue. Videos without readable captions are saved without a transcript, and Shorts are skipped.

Where are the files stored? On your machine, in videos/<channel>/<video-id>.md.

Which AI model does it use? google/gemini-2.5-flash by default through OpenRouter. You can change it with an environment variable or in the config file.

Is it okay to publish an AI blog made from someone else's video? Treat the output as a draft. Add your own analysis, credit and link the creator, embed the original video, and avoid mass-publishing near-copies. Google's spam policies target scaled, low-value content regardless of how it was made.

Conclusion

YouTube already contains the best explanations in software, AI, and startups. The bottleneck is turning that spoken knowledge into something you can search, link, and learn from. Tubelog automates the boring part, fetching and storing transcripts as Markdown, and gives you an AI first draft when you want one. The judgment, structure, and original insight are still yours to add.

Try it: clone Tubelog on GitHub, archive one channel you follow, and publish one good post from it this week. If you want more tools like this, the monthly iHateReading Magazine curates blogs, tools, directories, and trending repos by section, and the newsletter sends a weekly digest.

External Links (13)

Explore topics

More in

Weekly letter

Subscribe to iHateReading

Our once-a-week newsletter on programming, jobs, AI, and building products online.

Our once a week newsletter on Programming, Jobs, AI, and Business