📚 ScalerGPT#
Level 2 · 2 containers. Open a topic to see the idea, the request path and the function calls behind the demo, then read the complete Python source file by file.
How It Works#
The idea behind the demo, the request it sends and the function calls that answer it.
Concept#
This demo runs two containers: a FastAPI backend and a ChromaDB vector database. It uses Retrieval-Augmented Generation (RAG). The user's query is first embedded and used to search ChromaDB for relevant notes. The matched notes are stuffed into the prompt, and GPT-4o-mini generates an answer using only that context.
Theory & Concepts#
The ScalerGPT demo showcases a multi-container architecture for Retrieval-Augmented Generation (RAG). Standard Large Language Models (LLMs) like GPT-4o-mini are great at general knowledge, but they don't know your specific private study notes and might hallucinate answers. ScalerGPT solves this by giving the LLM the exact information it needs to answer your question.
To achieve this, the demo uses two separate Docker containers working together:
- The Vector Database Container (ChromaDB): This container stores your ingested notes as "embeddings"βmathematical representations of the text's meaning. When you ask a question like "what is a container?", the system converts your question into a vector and asks ChromaDB to find the notes with the most similar meaning.
- The ScalerGPT Backend Container: This FastAPI service coordinates the RAG process. First, it hits the ChromaDB container to retrieve the top 3 most relevant chunks of your notes. Then, it takes those chunks and stuffs them into a "Context" section of the prompt.
- Context Injection: Instead of just asking the LLM the question directly, the backend gives it strict system instructions: "Answer the user's question using ONLY the context below." The LLM synthesizes the final answer based purely on the retrieved notes, ensuring accuracy and avoiding hallucinations.
Request flow#
Code flow#
scalergpt_ask"] B -->|proxy_request| C["scalergpt:8000
POST /ask"] C -->|query| D["chroma container
collection.query"] D -->|relevant chunks| C C -->|context + query| E["OpenAI API
gpt-4o-mini"] E -->|answer text| C C -->|JSON result| B B -->|JSON result| A
Source Code#
Every Python file this demo runs, complete and unedited: the feature code first, then the shared Flask routes.
scaler-gpt/ingest.py#
Loads every .txt and .md file from docs/ into the Chroma vector database.
"""
Loads every .txt and .md file from docs/ into the Chroma vector database.
Run this AFTER the containers are up:
docker compose exec scalergpt python ingest.py
"""
import glob
import os
import sys
import time
import chromadb
from chromadb.utils import embedding_functions
# β validate the API key and configure the embedding model
API_KEY = os.getenv("OPENAI_API_KEY", "").strip()
if not API_KEY or API_KEY.startswith("sk-paste"):
sys.exit("[ingest] OPENAI_API_KEY is missing. Put a real key in .env and recreate.")
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
api_key=API_KEY,
model_name="text-embedding-3-small",
)
# β‘ connect to Chroma, retrying until the service is ready
host = os.getenv("CHROMA_HOST", "localhost")
port = int(os.getenv("CHROMA_PORT", "8000"))
chroma = None
for attempt in range(10):
try:
chroma = chromadb.HttpClient(host=host, port=port)
chroma.heartbeat()
break
except Exception:
print(f"[ingest] waiting for chroma ({attempt + 1}/10)...", flush=True)
time.sleep(2)
if chroma is None:
sys.exit(f"[ingest] Could not reach chroma at {host}:{port}")
# β’ open the notes collection and find source files to ingest
collection = chroma.get_or_create_collection(name="notes", embedding_function=openai_ef)
files = sorted(glob.glob("docs/*.txt") + glob.glob("docs/*.md"))
if not files:
sys.exit("No files found in docs/ β add some .txt or .md notes first.")
# β£ prepare containers for every chunk, document, and source label
ids, documents, metadatas = [], [], []
for path in files:
# β€ read each file and split it into useful paragraph chunks
with open(path, "r", encoding="utf-8") as f:
text = f.read().strip()
# Simple chunking: split on blank lines, keep chunks of reasonable size
chunks = [c.strip() for c in text.split("\n\n") if len(c.strip()) > 40]
# β₯ attach stable ids and source metadata to every chunk
for i, chunk in enumerate(chunks):
ids.append(f"{os.path.basename(path)}::{i}")
documents.append(chunk)
metadatas.append({"source": os.path.basename(path)})
# β¦ stop if the files did not produce any useful chunks
if not documents:
sys.exit("Files found but no usable chunks. Make paragraphs a bit longer.")
# β§ upsert chunks so re-running this script refreshes the index safely
# upsert = add or overwrite, so re-running this script is safe
collection.upsert(ids=ids, documents=documents, metadatas=metadatas)
# β¨ report how many chunks were indexed
print(f"Ingested {len(documents)} chunks from {len(files)} file(s) β
")
print(f"Collection now holds {collection.count()} chunks total.")
scaler-gpt/app.py#
ScalerGPT - A RAG chatbot over your own notes.
"""
ScalerGPT - A RAG chatbot over your own notes.
Architecture:
[ you ] --> app (FastAPI, this file) --> chroma (vector DB, separate container)
|
+--> OpenAI API (embeddings + chat completion)
"""
import os
import sys
import time
import chromadb
from chromadb.utils import embedding_functions
from fastapi import FastAPI, HTTPException
from openai import OpenAI
from pydantic import BaseModel
app = FastAPI(title="ScalerGPT", description="RAG over your course notes")
# --- Fail loudly and clearly if the API key is missing ---------------------
API_KEY = os.getenv("OPENAI_API_KEY", "").strip()
if not API_KEY or API_KEY.startswith("sk-paste"):
sys.exit(
"\n[ScalerGPT] OPENAI_API_KEY is missing.\n"
" Fix: cp .env.example .env, put your real key in it, then\n"
" docker compose up -d --force-recreate\n"
)
llm = OpenAI(api_key=API_KEY)
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
api_key=API_KEY,
model_name="text-embedding-3-small",
)
CHROMA_HOST = os.getenv("CHROMA_HOST", "localhost")
CHROMA_PORT = int(os.getenv("CHROMA_PORT", "8000"))
def connect_to_chroma(retries: int = 30, delay: int = 2):
"""
Chroma's container takes a few seconds to boot. `depends_on` only waits for
the container to START, not to be READY - so we retry here instead of
crashing on the first refused connection.
"""
# β try to create a Chroma client and prove it is ready
for attempt in range(1, retries + 1):
try:
client = chromadb.HttpClient(host=CHROMA_HOST, port=CHROMA_PORT)
client.heartbeat()
print(f"[ScalerGPT] Connected to chroma at {CHROMA_HOST}:{CHROMA_PORT}")
return client
except Exception as e:
# β‘ wait before retrying while Chroma finishes booting
print(
f"[ScalerGPT] Waiting for chroma "
f"({attempt}/{retries}): {type(e).__name__}",
flush=True,
)
time.sleep(delay)
# β’ stop the service with a clear message after all retries fail
sys.exit(f"\n[ScalerGPT] Could not reach chroma at {CHROMA_HOST}:{CHROMA_PORT}\n")
chroma = connect_to_chroma()
collection = chroma.get_or_create_collection(name="notes", embedding_function=openai_ef)
class Question(BaseModel):
query: str
@app.get("/")
def health():
return {
"status": "ScalerGPT is live π",
"docs_indexed": collection.count(),
"chroma_host": CHROMA_HOST,
"chroma_port": CHROMA_PORT,
}
@app.post("/ask")
def ask(q: Question):
# β reject questions until the vector index has documents
if collection.count() == 0:
raise HTTPException(
status_code=400,
detail="No documents indexed yet. Run: docker compose exec scalergpt python ingest.py",
)
# β‘ RETRIEVE - find the most relevant chunks from the vector DB
hits = collection.query(query_texts=[q.query], n_results=3)
documents = hits.get("documents") or [[]]
context = "\n\n---\n\n".join(documents[0])
# β’ AUGMENT - stuff that context into the prompt
system_prompt = (
"You are ScalerGPT, a helpful teaching assistant. "
"Answer the user's question using ONLY the context below. "
"If the context does not contain the answer, say you don't know.\n\n"
f"CONTEXT:\n{context}"
)
# β£ GENERATE - let the LLM write the final answer
resp = llm.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": q.query},
],
)
# β€ return the answer and how many source chunks were used
return {
"question": q.query,
"answer": resp.choices[0].message.content,
"sources_used": len(documents[0]),
}
app.py#
Flask server proxying browser requests to internal Docker demo services.
"""Flask server proxying browser requests to internal Docker demo services.
Architecture notes
------------------
- All routes are attached to a Blueprint (``bp``) instead of directly to
``app``. This lets us register the entire Blueprint under a runtime URL
prefix (``PATH_PREFIX``) without touching individual route strings.
- In local development PATH_PREFIX is empty, so routes are at "/",
"/quickbite/predict", etc. In production Nginx forwards ``/docker/...``
traffic to the container and PATH_PREFIX is set to "/docker".
- flask-cors adds ``Access-Control-Allow-Origin: *`` headers so the HTML
page can call the API even if it is served from a different origin during
development.
- Proxy routes forward browser requests to internal Docker services
(quickbite, scalergpt, deskbuddy-agent) using service-name networking.
"""
import os
from pathlib import Path
import requests as http_client
from flask import Blueprint, Flask, jsonify, request
from flask_cors import CORS
from rate_limiter import check_rate_limit
# ---------------------------------------------------------------------------
# Configuration
# ---------------------------------------------------------------------------
# PATH_PREFIX is set by the deployment environment (e.g. "/docker") so the app
# works correctly behind an Nginx location block. Locally it is an empty
# string, which mounts all routes at the root.
PATH_PREFIX = os.environ.get("PATH_PREFIX", "")
# app.py lives in src/python, while index.html, css/, and js/ live in src/.
STATIC_DIR = Path(__file__).resolve().parents[1]
app = Flask(__name__, static_folder=str(STATIC_DIR))
# Allow cross-origin requests from any origin. In production you would
# restrict this to the specific front-end domain.
CORS(app)
# A Blueprint groups related routes. We register it once at the bottom with
# the runtime PATH_PREFIX, avoiding any hardcoded path strings in the routes.
bp = Blueprint("main", __name__)
@bp.before_request
def enforce_rate_limit():
"""Enforce strict 10 requests per hour limit on all POST endpoints."""
# β only rate-limit write requests so page assets stay fast
if request.method == "POST":
# β‘ check the caller's hourly quota before proxying work
blocked, msg, retry_after = check_rate_limit(
request, max_requests=10, window_seconds=3600
)
# β’ return a 429 with Retry-After when the quota is used up
if blocked:
resp = jsonify({"error": msg})
resp.status_code = 429
resp.headers["Retry-After"] = str(retry_after)
return resp
# Internal service URLs β these use Docker Compose service names, never IPs.
QUICKBITE_URL = "http://quickbite:8000"
SCALERGPT_URL = "http://scalergpt:8000"
DESKBUDDY_URL = "http://deskbuddy-agent:9000"
# Timeout for proxy requests to example services (seconds).
PROXY_TIMEOUT = 30
# ---------------------------------------------------------------------------
# Helper
# ---------------------------------------------------------------------------
def proxy_request(method, url, json_body=None):
"""Forward a request to an internal service and return its JSON response.
Returns a tuple of (response_dict, http_status_code). On connection
errors, returns a helpful error message instead of crashing.
"""
# β forward the request to the selected internal service
try:
if method == "GET":
resp = http_client.get(url, timeout=PROXY_TIMEOUT)
else:
resp = http_client.post(url, json=json_body, timeout=PROXY_TIMEOUT)
# β‘ pass through the service JSON and HTTP status code
return resp.json(), resp.status_code
except http_client.ConnectionError:
# β’ turn connection failures into a helpful service-start message
service = url.split("//")[1].split(":")[0]
return {
"error": f"Service '{service}' is not running. "
f"Start it with: docker compose up {service}"
}, 503
except Exception as e:
# β£ return unexpected proxy failures as JSON instead of crashing
return {"error": str(e)}, 500
# ---------------------------------------------------------------------------
# Routes β Static files
# ---------------------------------------------------------------------------
@bp.route("/")
def index():
"""Serve index.html, injecting the correct API base URL for the environment."""
# β read the static homepage template from the shared src folder
with open(os.path.join(app.static_folder, "index.html"), encoding="utf-8") as f:
html = f.read()
# β‘ inject the runtime API prefix so browser calls hit this gateway
html = html.replace('data-api-base=""', f'data-api-base="{PATH_PREFIX}"')
# β’ return the rendered HTML with the correct MIME type
return app.response_class(html, mimetype="text/html")
@bp.route("/css/<path:filename>")
def css(filename):
"""Serve stylesheets from the src/css directory."""
return app.send_static_file(os.path.join("css", filename))
@bp.route("/js/<path:filename>")
def js(filename):
"""Serve scripts from the src/js directory."""
return app.send_static_file(os.path.join("js", filename))
@bp.route("/info/<path:filename>")
def info(filename):
"""Serve the "how this demo works" explainer pages from src/info."""
return app.send_static_file(os.path.join("info", filename))
# ---------------------------------------------------------------------------
# Routes β QuickBite ETA (Level 1, keyless)
# ---------------------------------------------------------------------------
@bp.route("/quickbite/predict", methods=["POST"])
def quickbite_predict():
"""Proxy ETA prediction to the QuickBite FastAPI service."""
# β parse the browser's order JSON
body = request.get_json(force=True)
# β‘ proxy the order to the QuickBite prediction service
data, status = proxy_request("POST", f"{QUICKBITE_URL}/predict", body)
# β’ return the service JSON and status code unchanged
return jsonify(data), status
@bp.route("/quickbite/status")
def quickbite_status():
"""Check if QuickBite service is running."""
# β ask QuickBite for its health payload
data, status = proxy_request("GET", f"{QUICKBITE_URL}/")
# β‘ return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Routes β ScalerGPT (Level 2, needs OPENAI_API_KEY)
# ---------------------------------------------------------------------------
@bp.route("/scalergpt/ask", methods=["POST"])
def scalergpt_ask():
"""Proxy RAG question to the ScalerGPT FastAPI service."""
# β parse the browser's question JSON
body = request.get_json(force=True)
# β‘ proxy the question to the ScalerGPT RAG service
data, status = proxy_request("POST", f"{SCALERGPT_URL}/ask", body)
# β’ return the service JSON and status code unchanged
return jsonify(data), status
@bp.route("/scalergpt/status")
def scalergpt_status():
"""Check if ScalerGPT service is running and how many docs are indexed."""
# β ask ScalerGPT for its health and index summary
data, status = proxy_request("GET", f"{SCALERGPT_URL}/")
# β‘ return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Routes β DeskBuddy (Level 3, needs OPENAI_API_KEY)
# ---------------------------------------------------------------------------
@bp.route("/deskbuddy/chat", methods=["POST"])
def deskbuddy_chat():
"""Proxy chat message to the DeskBuddy agent service."""
# β parse the browser's chat JSON
body = request.get_json(force=True)
# β‘ proxy the message to the DeskBuddy agent loop
data, status = proxy_request("POST", f"{DESKBUDDY_URL}/chat", body)
# β’ return the agent JSON and status code unchanged
return jsonify(data), status
@bp.route("/deskbuddy/status")
def deskbuddy_status():
"""Check if DeskBuddy agent service is running."""
# β ask DeskBuddy for its health payload
data, status = proxy_request("GET", f"{DESKBUDDY_URL}/")
# β‘ return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Blueprint registration + server entry point
# ---------------------------------------------------------------------------
app.register_blueprint(bp, url_prefix=PATH_PREFIX)
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)