Briefing

Synthadoc v0.3.0: Ingest YouTube Videos and Web Search into a Structured Wiki

ai-dev
by Paul Chen ·

Patch: integrate Synthadoc v0.3.0 into your knowledge base pipeline to ingest YouTube videos and web search results with timestamped transcripts and cross‑references.

What to do now

Patch: add Synthadoc v0.3.0 to your repo, run `synthadoc ingest` for your knowledge sources, and verify cross‑references and timestamps.

Summary

Synthadoc v0.3.0 (released 2026‑05‑03) expands its ingestion pipeline to include YouTube videos and live web search results, in addition to the previously supported PDFs, Word, XLSX, CSV, TXT, images, and PowerPoint files. The new YouTube workflow fetches the caption track, chunks the transcript with embedded [MM:SS] timestamps, generates an executive summary, and creates a structured Markdown wiki page that cross‑references existing pages. Web search ingestion now performs a multi‑query fan‑out, synthesizing each result into a wiki page and building cross‑references across all new pages. All source types produce the same output format: a Markdown page with frontmatter, wikilinks, and traceable source references. The pipeline also supports Vision LLM extraction for images and OCR for PDFs, and the new version includes a command‑line interface for bulk ingestion.

The release adds several new features: YouTube caption ingestion, timestamped transcripts, executive summaries, cross‑linking, multi‑query web search, support for images via Vision LLM, and a unified output format. The tool now requires Python 3.10+, the bleak BLE library for future extensions, and the facturx library for optional PDF embedding.

Overall, v0.3.0 turns Synthadoc into a full‑stack knowledge ingestion engine that can automatically transform any media source into a searchable, cross‑referenced wiki.

Key changes

  • Ingests YouTube captions with embedded [MM:SS] timestamps
  • Generates an executive summary for each video
  • Cross‑references new wiki pages to existing ones
  • Multi‑query fan‑out web search synthesizes each result into a wiki page
  • Supports PDFs, Word, XLSX, CSV, TXT, images via Vision LLM, and PowerPoint
  • All source types output a unified Markdown wiki page with frontmatter
  • Adds command‑line interface for bulk ingestion
  • Requires Python 3.10+, bleak, and facturx for optional PDF embedding

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting