A Multi-Agent, Source-Grounded Educational AI System for Adaptive and Cross-Domain Learning
EduPilot is an intelligent course assistant built for Indiana University students. It answers questions across four graduate courses — Applied Machine Learning (AML), Applied Database Technologies (ADT), Statistics (STAT), and Large Language Models (LLM) — using a seven-stage multi-agent RAG pipeline that retrieves directly from course lecture slides and materials, cites every claim, verifies its own answers, and refuses to hallucinate.
Built by: Akshar Patel · Khushi Shah
Institution: Indiana University Bloomington
- Overview
- Key Features
- System Architecture
- Screenshots
- Tech Stack
- Project Structure
- Setup and Installation
- Running the Application
- Evaluation Suite
- Configuration
- API Reference
- Authors
EduPilot solves a core problem in educational AI: LLMs hallucinate. Generic chatbots answer confidently with no connection to actual course materials, making them unreliable for exam prep or homework help.
EduPilot's approach:
- Every answer is grounded in the course knowledge base — retrieved chunks from real lecture slides, not GPT's parametric memory.
- Every claim is cited with a
[Source N]marker linking back to the exact lecture and page number. - Every answer is verified — a two-pass verifier checks quality (≥ 0.75) and coverage (≥ 0.70), and rewrites the answer if either threshold is missed.
- Cross-domain queries are decomposed into sub-questions, answered independently per domain, and synthesized into a unified response.
- Out-of-domain questions are politely refused rather than answered with hallucinated content.
| Feature | Description |
|---|---|
| Multi-agent pipeline | 7 independent agents: Router → Splitter → Retriever → Reranker → Generator → Synthesizer → Verifier |
| Hybrid retrieval | Reciprocal Rank Fusion (60% semantic + 40% BM25) outperforms unimodal retrieval by 8–14% |
| Source citations | Every answer includes [Source N] markers with lecture name and page number |
| Two-pass verification | Self-grading on quality and coverage; targeted rewrite if thresholds not met |
| Cross-domain synthesis | Automatically detects multi-domain queries and retrieves from each domain separately |
| Out-of-domain guard | Hard refusal for questions outside AML, ADT, STAT, LLM scope |
| Self Study Mode | Upload any personal documents (PDF, TXT, DOCX) and chat with them privately |
| Evaluation dashboard | Built-in UI to run the 10-query test suite and see per-case metrics |
| Model selector | Switch between Groq (Llama 3.3 70B, Llama 3.1 8B) and Gemini fallback at runtime |
| Debug panel | Real-time view of retrieved chunks, reranking scores, and verification reasoning |
EduPilot processes every query through a seven-stage pipeline:
Student Query
│
▼
┌─────────────────────────────────┐
│ Stage 1 — ROUTER │ Classifies intent (single / multi / OOD)
│ Intent & domain detection │ Keyword fallback if LLM API fails
│ + clarification guard │
└──────────────┬──────────────────┘
│
┌─────────┴──────────┐
│ single-domain │ multi-domain
▼ ▼
┌──────────┐ ┌──────────────────────────┐
│ Stage 2 │ │ Stage 2 — QUERY │
│ (bypass) │ │ SPLITTER │ Decomposes into N sub-questions
└────┬─────┘ └──────────┬───────────────┘
│ │ (one branch per domain)
▼ ▼
┌─────────────────────────────────┐
│ Stage 3 — HYBRID RETRIEVER │ Pinecone (semantic) + BM25 (keyword)
│ RRF(c) = Σ w_r / (k + rank_r) │ k=60, w_sem=0.60, w_bm25=0.40
│ all-MiniLM-L6-v2 (384-dim) │ Top-K = 8 candidates
└──────────────┬──────────────────┘
│
▼
┌─────────────────────────────────┐
│ Stage 4 — RERANKER │ Confidence-thresholded cross-encoder
│ Filters to Top-K = 5 chunks │ Threshold = 0.20
└──────────────┬──────────────────┘
│
▼
┌─────────────────────────────────┐
│ Stage 5 — DOMAIN AGENT(S) │ One LLM call per domain
│ Groq Llama 3.3 70B │ Prompt: retrieved context + citations
│ Gemini fallback (auto) │
└──────────────┬──────────────────┘
│
▼
┌─────────────────────────────────┐
│ Stage 6 — SYNTHESIZER │ Merges multi-domain answers
│ (single-domain: pass-through) │ into one coherent response
└──────────────┬──────────────────┘
│
▼
┌─────────────────────────────────┐
│ Stage 7 — VERIFIER │ Scores quality (≥ 0.75) + coverage (≥ 0.70)
│ Two-pass self-grading │ Targeted rewrite if below threshold
│ 4.6× token efficiency gain │
└──────────────┬──────────────────┘
│
▼
Final Answer
(cited, verified, grounded)
| Domain | Code | Colour | Coverage |
|---|---|---|---|
| Applied Machine Learning | AML | Green | Supervised/unsupervised learning, deep learning, optimisation |
| Applied Database Technologies | ADT | Blue | SQL, NoSQL, normalisation, transactions, data warehousing |
| Statistics | STAT | Orange | Probability, hypothesis testing, regression, Bayesian inference |
| Large Language Models | LLM | Purple | Transformers, attention, fine-tuning, RAG, prompt engineering |
The full EduPilot UI — conversation history in the left sidebar, model and retrieval settings below it, suggested questions at the top, and the main answer pane. Domain tags (AML, ADT, STAT, LLM) badge every response.
EduPilot answers a cross-domain question about Backpropagation and LLM Training with a 96% quality score. The response is structured into Overview, Core Concepts, Algorithm definition, and Mechanism sections — all grounded in retrieved course material.
A query spanning three domains ("Hypothesis Testing, Supervised ML, and RAG optimisation techniques") is automatically decomposed. Each part is retrieved independently, then synthesized into one unified answer. Domain badges show which knowledge base each section came from.
Every answer includes an expandable citations panel showing exactly which lecture slide and page each source came from. Sources span multiple domains (STAT, LLM, AML) in a single response and are downloadable as PDFs.
When a question falls outside the four supported courses (e.g., "What are the important components in finance?"), EduPilot refuses to answer rather than hallucinate, and clearly lists the four supported domains.
Instructors or students can extend any domain knowledge base by attaching PDF, TXT, or DOCX files through the drag-and-drop upload modal. Files are indexed into Pinecone and BM25 automatically.
Self Study Mode lets students upload any personal documents — notes, textbooks, past papers — and chat with them privately. This is completely separate from the course knowledge base and does not affect other users.
An active "ML Interview" study session with two uploaded PDFs (38 chunks). EduPilot answers a question about Multi-Head Attention with inline [Source N] citations drawn exclusively from the uploaded documents.
The built-in Evaluation tab shows live results across all 10 test cases — 100% intent accuracy, 100% domain accuracy, retrieval hit rate 0.49, mean quality 0.84, and citation accuracy 1.00 with the primary model (Llama-3.3-70B). The suite can be run from this tab or via python3 run_eval.py.
| Layer | Technology |
|---|---|
| LLM | Groq (Llama 3.3 70B Versatile, Llama 3.1 8B Instant) · Gemini 2.0 Flash (fallback) |
| Embeddings | all-MiniLM-L6-v2 (384-dim, sentence-transformers) |
| Vector Store | Pinecone Serverless (AWS us-east-1) |
| Keyword Search | rank-bm25 (in-memory BM25 index, rebuilt from SQLite on startup) |
| Reranking | Confidence-thresholded score filtering (threshold = 0.20) |
| Backend | FastAPI + Uvicorn (async) |
| Frontend | Streamlit (app.py) |
| Database | SQLite (conversation history, BM25 chunk cache) |
| Document Parsing | PyMuPDF (PDF) · python-docx (DOCX) · plain text |
| Environment | Python 3.10+ · python-dotenv |
EduPilot/
├── main.py # FastAPI app + _run_pipeline orchestrator
├── app.py # Streamlit frontend
├── config.py # Central config (models, domains, thresholds)
├── router.py # Stage 1: intent classification + domain routing
├── query_splitter.py # Stage 2: multi-domain query decomposition
├── retriever.py # Stage 3: Pinecone + BM25 hybrid retrieval
├── reranker.py # Stage 4: confidence-thresholded reranking
├── synthesizer.py # Stage 6: multi-domain answer synthesis
├── verifier.py # Stage 7: two-pass quality verification
├── prompts.py # All seven LLM prompt templates
├── utils.py # LLM caller, document chunking, shared types
├── database.py # SQLite session + message storage
├── evaluation.py # 10-query evaluation suite + 8 metrics
├── model_comparison.py # 3-model comparison runner (saves one CSV per model)
├── run_eval.py # Standalone evaluation runner script
├── self_study_retriever.py # Private document retrieval for Self Study Mode
│
├── knowledge_base/
│ ├── aml/ # Applied Machine Learning lecture slides
│ ├── adt/ # Applied Database Technologies materials
│ ├── stats/ # Statistics lecture notes
│ ├── llm/ # Large Language Models course materials
│ └── devops/ # DevOps supplementary materials
│
├── screenshots/ # UI screenshots (used in this README)
├── requirements.txt
└── .env # API keys (not committed)
- Python 3.10 or higher
- A Groq API key (free tier available)
- A Pinecone API key with a serverless index named
edupilot - (Optional) A Google Gemini API key for fallback
git clone https://github.com/Akshar106/EduPilot-A-Multi-Agent-Source-Grounded-Educational-AI-System-for-Adaptive-and-Cross-Domain-Learning.git
cd EduPilot-A-Multi-Agent-Source-Grounded-Educational-AI-System-for-Adaptive-and-Cross-Domain-Learningpython3 -m venv venv
source venv/bin/activate # macOS / Linux
# venv\Scripts\activate # Windowspip install -r requirements.txtCreate a .env file in the project root:
# Required
GROQ_API_KEY=your_groq_api_key_here
PINECONE_API_KEY=your_pinecone_api_key_here
PINECONE_INDEX_NAME=edupilot
PINECONE_CLOUD=aws
PINECONE_REGION=us-east-1
# Optional — Gemini fallback when Groq quota is exhausted
GEMINI_API_KEY=your_gemini_api_key_here
# Database
SQLITE_DB_PATH=edupilot.dbPlace PDF, TXT, or DOCX files into the appropriate subdirectory:
knowledge_base/aml/ ← AML lecture slides
knowledge_base/adt/ ← ADT materials
knowledge_base/stats/ ← Statistics notes
knowledge_base/llm/ ← LLM course materials
Documents are indexed into Pinecone and the BM25 cache automatically on first startup.
uvicorn main:app --host 0.0.0.0 --port 8000 --reloadAPI available at http://localhost:8000.
Interactive docs: http://localhost:8000/docs
streamlit run app.pyUI opens at http://localhost:8501.
EduPilot ships with a 10-query evaluation suite covering all pipeline layers across three categories:
| Category | Queries | Description |
|---|---|---|
| Single-domain | 6 | Factual and conceptual queries spanning AML, ADT, STAT, and LLM domains |
| Multi-domain | 2 | Cross-domain queries requiring two-namespace synthesis |
| Adversarial | 2 | Fabricated concepts and false-premise queries stressing hallucination resistance |
| Total | 10 |
| Category | N | Intent | Hit Rate | Quality |
|---|---|---|---|---|
| Single-domain | 6 | 1.00 | 0.53 | 0.91 |
| Multi-domain | 2 | 1.00 | 0.50 | 0.70 |
| Adversarial | 2 | 1.00 | 0.20 | 0.70 |
| Overall | 10 | 1.00 | 0.49 | 0.84 |
Citation accuracy: 1.00 across all generating queries. No verifier revisions triggered.
| Metric | What it measures |
|---|---|
| Intent Match | Router correctly classified single vs. multi intent |
| Domain Match | Router routed to the correct domain(s) |
| Retrieval Hit Rate | Fraction of expected keywords found in retrieved chunks |
| Faithfulness | LLM-judged grounding of answer in retrieved evidence |
| Citation Accuracy | [Source N] markers match their referenced chunks |
| Quality Score | Verifier rubric score (0.95–1.00 Exceptional, 0.70–0.84 Adequate) |
| Coverage Score | Sub-topics addressed relative to expected behavior |
| Latency (ms) | End-to-end wall-clock time |
| Metric | Llama-3.3-70B | Llama-3.1-8B | Gemini-2.5-Flash |
|---|---|---|---|
| Quality score | 0.842 | 0.830 | 0.847 |
| Citation accuracy | 1.000 | 0.270 | 1.000 |
| Faithfulness | 0.662 | 0.725 | 0.450 |
| Avg latency (ms) | 13,785 | 15,073 | 15,117 |
Key finding: a 3.7× citation-accuracy gap between Llama-70B and Llama-8B on identical inputs — purely from model choice.
# Run all 10 test cases
python3 run_eval.py
# Run the 3-model comparison (saves one CSV per model)
python3 model_comparison_eval.py
# Save results to a timestamped file
python3 run_eval.py --out results_$(date +%Y%m%d).jsonResults are printed to stdout and saved as JSON/CSV. The Evaluation tab in the Streamlit UI provides the same functionality with a live visual dashboard (see screenshot above).
All tunable parameters live in config.py:
# Retrieval
DEFAULT_TOP_K = 8 # Candidates from Pinecone + BM25
DEFAULT_RERANK_TOP_K = 5 # Chunks passed to LLM after reranking
DEFAULT_CONFIDENCE_THRESHOLD = 0.20
# Hybrid RRF weights
SEMANTIC_WEIGHT = 0.60
BM25_WEIGHT = 0.40
RRF_K = 60 # RRF(c) = Σ w / (k + rank)
# Verification thresholds
QUALITY_THRESHOLD = 0.75 # Min quality score before rewrite
COVERAGE_THRESHOLD = 0.70 # Min coverage score before rewrite
# Chunking
CHUNK_SIZE = 800
CHUNK_OVERLAP = 150
# Models
DEFAULT_MODEL = "llama-3.3-70b-versatile" # Groq
VERIFY_MODEL = "llama-3.3-70b-versatile"
EMBEDDING_MODEL = "all-MiniLM-L6-v2"Send a query through the full pipeline.
Request:
{
"query": "What is the bias-variance tradeoff?",
"session_id": "abc123",
"model": "llama-3.3-70b-versatile",
"top_k": 8,
"rerank_top_k": 5,
"confidence_threshold": 0.20,
"enable_verification": true
}Response:
{
"final_answer": "The bias-variance tradeoff describes...",
"intent_type": "single",
"detected_domains": ["AML"],
"quality_score": 0.92,
"sources": [
{ "source_num": 1, "citation_label": "AML · Lec3 p.12", "text": "..." }
],
"needs_clarification": false,
"is_course_related": true
}Run a single named test case from the evaluation suite.
Run all 10 test cases and return aggregate statistics.
Upload a document (PDF / TXT / DOCX) to a domain knowledge base.
Upload a document to a private Self Study session.
| Name | Contributions | |
|---|---|---|
| Akshar Patel | akspate@iu.edu | System architecture, hybrid retrieval pipeline (Pinecone + BM25 + RRF), query router with keyword fallback, query splitter, SQLite database layer, FastAPI async backend, confidence-thresholded reranker, evaluation suite construction, report writing |
| Khushi Shah | khusshah@iu.edu | Domain agent prompt engineering, cross-domain synthesizer, two-pass verifier with targeted revision, Self Study module, Streamlit debug UI, all seven prompt templates, evaluation framework (10 queries, 8 metrics, three categories), 3-model comparison analysis, primary report writing |
Indiana University Bloomington · Luddy School of Informatics, Computing, and Engineering
EduPilot — A TA that never sleeps.








