Skip to content
R&D prototypeRAG · Hallucination guard · ZAKA capstone

YTBOT — Answers from YouTube, with verified citations

Ask any question; YTBOT finds relevant YouTube videos, retrieves transcript passages, answers from them — and rejects any citation it cannot find verbatim in the transcript.

Asking a model to cite is not enough. Verify the citation, or don't answer.

01Context

In late 2023, chatbots answered from training data frozen at a cut-off date and favoured sounding right over being right. YouTube is current and human-made, but finding one answer means watching whole videos.

02Problem

Answer any question from YouTube without the user choosing a video — with exact, clickable citations, and an honest “not found” when the videos don't say.

03Constraints

  • Model context limits (4k tokens on GPT-3.5 at the time)
  • Cost per call grows with tokens
  • Existing open-source tools required picking the video and gave no citations

04My role

What I personally built

  • Co-designed and co-built the full pipeline with Ali Faraj — the individual split is not documented, so no part is claimed as solely mine

What others owned

  • Ali Faraj — co-author

05Architecture

Architecture explorer

The business flow — what happens, in plain words.

01 / 05

Question — Anything, in plain words.

06Key decisions

  1. Verify every citation against the transcript before showing it

    Because
    A fabricated citation is replaced by “I could not find the answer from the YouTube videos”.
    Trade-off
    Some correct paraphrased answers are rejected too.
  2. Keep a single dense retriever over the hybrid pipeline

    Because
    It found relevant passages even when the question was phrased differently from the source.
    Trade-off
    Less robust to exact-term queries than keyword retrieval.

07System

An LLM turns the question into search keywords; the YouTube Data API returns the five most relevant videos; transcripts are split into ~3,000-character chunks, embedded and stored in ChromaDB; the four most similar chunks become the context; the model answers with structured output (Pydantic) separating answer from citation; every citation is then searched for in the transcript before it is shown, with a timestamped link.

08Challenges

  • Models invent citations even when told not to
  • Choosing between hybrid retrieval (sparse + dense, merged) and a single dense retriever

09Outcome

A working Streamlit app with search, answer, citation and video sections. Both retrieval designs were tested; the simpler single dense retriever was kept because it was good enough without the added complexity.

10What I learned

Grounding is an engineering control, not a prompt. That idea runs straight through to this portfolio's own assistant.

11Stack

  • Python
  • OpenAI GPT-3.5
  • OpenAI embeddings
  • ChromaDB
  • LangChain
  • YouTube Data API v3
  • Streamlit
  • Pydantic

12Evidence