🤖 DeskBuddy#
Level 3 · 3 containers. Open a topic to see the idea, the request path and the function calls behind the demo, then read the complete Python source file by file.
How It Works#
The idea behind the demo, the request it sends and the function calls that answer it.
Concept#
A multi-container agentic system with three containers: the deskbuddy-agent (FastAPI + LLM logic), a tools service (for executing functions like calculator/datetime), and redis for conversation memory. The LLM decides when to call a tool, the agent pauses to call the tools service, and the result is fed back into the prompt in a loop until the final answer is reached.
Theory & Concepts#
The DeskBuddy demo illustrates a production-ready Agentic Architecture distributed across multiple Docker containers. At its core, an agent uses a Large Language Model (LLM) not just to generate text, but to iteratively reason and take action—specifically, by deciding when to call external functions (tools). In this demo, when you ask DeskBuddy to "calculate 23 * 47" or "what time is it", the LLM realizes it needs external help and emits a structured JSON request rather than a direct answer.
To make this secure and scalable, the demo adopts a Microservices approach with three isolated containers:
- The Agent Container: This FastAPI service runs the core loop. It receives your chat, asks the GPT-4o-mini model for a response, and handles the conditional logic. If the LLM requests a tool call, the agent pauses to dispatch the request to the Tools Container, then feeds the result back to the LLM to form a final answer.
- The Tools Container: Running tools in the same process as your agent can be dangerous, especially with features like math evaluation. The DeskBuddy architecture isolates the
/calculatorand/datetimeendpoints in a separate, sandboxed service. The agent communicates with this container over a private Docker network, minimizing security risks. - The Redis Container (Memory): LLMs are stateless by nature. To remember your past messages, the DeskBuddy agent loads and saves the conversation history from a Redis database using your session ID. By externalizing state, the agent container itself remains stateless and resilient.
Request flow#
Code flow#
deskbuddy_chat"] B -->|proxy_request| C["deskbuddy-agent:9000
POST /chat"] C -->|load history| D["Redis
history:session_id"] D -->|history| C C -->|messages + tools| E["OpenAI API
gpt-4o-mini"] E -.->|tool_calls| C C -.->|call_tool| F["tools:7000
/calculator or /datetime"] F -.->|result| C C -.->|append tool result| E E -->|final answer| C C -->|save history| D C -->|answer| B B -->|answer| A
Source Code#
Every Python file this demo runs, complete and unedited: the feature code first, then the shared Flask routes.
desk-buddy/agent/app.py#
DeskBuddy Agent - an LLM with hands.
"""
DeskBuddy Agent - an LLM with hands.
The agent loop:
1. Show the LLM the conversation + a menu of available tools
2. If the LLM answers in words -> done
3. If the LLM says "run tool X with these arguments" -> we call the
tools service over HTTP, paste the result back, and loop again
4. Max 5 laps (a safety fuse so a confused model can't loop forever)
Memory:
Conversation history is stored in Redis per session_id, so the agent
remembers context across requests - and across container restarts,
because Redis has a volume.
"""
import json
import os
import sys
import httpx
import redis
from fastapi import FastAPI
from openai import OpenAI
from pydantic import BaseModel
app = FastAPI(title="DeskBuddy Agent")
# --- Fail loudly and clearly if the key is missing --------------------------
API_KEY = os.getenv("OPENAI_API_KEY", "").strip()
if not API_KEY or API_KEY.startswith("sk-paste"):
sys.exit(
"\n[DeskBuddy] OPENAI_API_KEY is missing.\n"
" Fix: cp .env.example .env, put your real key in it, then\n"
" docker compose up -d --force-recreate\n"
)
llm = OpenAI(api_key=API_KEY)
# --- Service addresses: SERVICE NAMES, not IPs (Compose networking) ---------
# redis-py connects lazily (only on first command), so no retry loop needed:
# by the time the first /chat request arrives, Redis is long since ready.
r = redis.Redis(
host=os.getenv("REDIS_HOST", "redis"),
port=6379,
decode_responses=True,
)
TOOLS_URL = os.getenv("TOOLS_URL", "http://tools:7000")
# --- The tool menu we show the LLM ------------------------------------------
# Each entry describes one tool: its name, what it does, and what inputs
# it expects. The LLM reads these descriptions to decide when to use them.
TOOL_DEFS = [
{
"type": "function",
"function": {
"name": "calculator",
"description": "Evaluate a math expression, e.g. '23*47' or '(100-8)/4'",
"parameters": {
"type": "object",
"properties": {"expression": {"type": "string"}},
"required": ["expression"],
},
},
},
{
"type": "function",
"function": {
"name": "get_datetime",
"description": "Get the current date and time",
"parameters": {"type": "object", "properties": {}},
},
},
]
def call_tool(name: str, args: dict) -> dict:
"""Actually execute a tool by calling the tools microservice over HTTP."""
try:
# ① send calculator requests to the calculator endpoint
if name == "calculator":
return httpx.post(f"{TOOLS_URL}/calculator", json=args, timeout=10).json()
# ② send datetime requests to the datetime endpoint
if name == "get_datetime":
return httpx.get(f"{TOOLS_URL}/datetime", timeout=10).json()
# ③ report unknown tool names instead of guessing
return {"error": f"unknown tool: {name}"}
except Exception as e:
# ④ turn tool-service failures into JSON the agent can read
return {"error": f"tool call failed: {e}"}
class Chat(BaseModel):
session_id: str
message: str
@app.get("/")
def health():
return {
"status": "DeskBuddy Agent is live 🤖",
"tools_url": TOOLS_URL,
"redis_host": os.getenv("REDIS_HOST", "redis"),
}
@app.post("/chat")
def chat(req: Chat):
# ① load this session's memory from Redis and append the new user message
key = f"history:{req.session_id}"
history = [json.loads(m) for m in r.lrange(key, 0, -1)]
history.append({"role": "user", "content": req.message})
# ② run the bounded agent loop: think, act, observe, repeat
msg = None
for _ in range(5): # safety fuse: max 5 laps
resp = llm.chat.completions.create(
model="gpt-4o-mini",
messages=history,
tools=TOOL_DEFS,
)
msg = resp.choices[0].message
# ③ stop looping when the model answers in normal text
if not msg.tool_calls:
break # the LLM answered in words - we're done
# ④ run one or more tools and append their results for the next lap
# The LLM asked to run one or more tools
history.append(msg.model_dump(exclude_none=True))
for tc in msg.tool_calls:
result = call_tool(tc.function.name, json.loads(tc.function.arguments))
history.append(
{
"role": "tool",
"tool_call_id": tc.id,
"content": json.dumps(result),
}
)
# ⑤ save memory back to Redis and return the final answer
history.append({"role": "assistant", "content": msg.content})
r.delete(key)
for m in history:
r.rpush(key, json.dumps(m))
return {"answer": msg.content}
desk-buddy/tools/app.py#
DeskBuddy Tools - a tiny microservice exposing two tools: POST /calculator -> evaluates a math expression GET /datetime -> returns the current date & time
"""
DeskBuddy Tools - a tiny microservice exposing two tools:
POST /calculator -> evaluates a math expression
GET /datetime -> returns the current date & time
Nothing AI about this file. It's a plain worker department.
The agent (a separate container) calls these over the private Docker network.
"""
import datetime
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI(title="DeskBuddy Tools")
class Calc(BaseModel):
expression: str
@app.get("/")
def health():
return {"status": "DeskBuddy Tools is live 🧰"}
@app.post("/calculator")
def calculator(c: Calc):
"""Evaluate a math expression like '23*47' or '(100-8)/4'."""
try:
# ① demo only — never use eval on untrusted input in production!
# The empty __builtins__ blocks access to dangerous functions,
# but a real system would use a proper math parser.
result = eval(c.expression, {"__builtins__": {}})
return {"result": result}
except Exception as e:
# ② return parser or evaluation errors as JSON for the agent
return {"error": str(e)}
@app.get("/datetime")
def now():
"""Return the current date and time in ISO format."""
return {"now": datetime.datetime.now().isoformat()}
app.py#
Flask server proxying browser requests to internal Docker demo services.
"""Flask server proxying browser requests to internal Docker demo services.
Architecture notes
------------------
- All routes are attached to a Blueprint (``bp``) instead of directly to
``app``. This lets us register the entire Blueprint under a runtime URL
prefix (``PATH_PREFIX``) without touching individual route strings.
- In local development PATH_PREFIX is empty, so routes are at "/",
"/quickbite/predict", etc. In production Nginx forwards ``/docker/...``
traffic to the container and PATH_PREFIX is set to "/docker".
- flask-cors adds ``Access-Control-Allow-Origin: *`` headers so the HTML
page can call the API even if it is served from a different origin during
development.
- Proxy routes forward browser requests to internal Docker services
(quickbite, scalergpt, deskbuddy-agent) using service-name networking.
"""
import os
from pathlib import Path
import requests as http_client
from flask import Blueprint, Flask, jsonify, request
from flask_cors import CORS
from rate_limiter import check_rate_limit
# ---------------------------------------------------------------------------
# Configuration
# ---------------------------------------------------------------------------
# PATH_PREFIX is set by the deployment environment (e.g. "/docker") so the app
# works correctly behind an Nginx location block. Locally it is an empty
# string, which mounts all routes at the root.
PATH_PREFIX = os.environ.get("PATH_PREFIX", "")
# app.py lives in src/python, while index.html, css/, and js/ live in src/.
STATIC_DIR = Path(__file__).resolve().parents[1]
app = Flask(__name__, static_folder=str(STATIC_DIR))
# Allow cross-origin requests from any origin. In production you would
# restrict this to the specific front-end domain.
CORS(app)
# A Blueprint groups related routes. We register it once at the bottom with
# the runtime PATH_PREFIX, avoiding any hardcoded path strings in the routes.
bp = Blueprint("main", __name__)
@bp.before_request
def enforce_rate_limit():
"""Enforce strict 10 requests per hour limit on all POST endpoints."""
# ① only rate-limit write requests so page assets stay fast
if request.method == "POST":
# ② check the caller's hourly quota before proxying work
blocked, msg, retry_after = check_rate_limit(
request, max_requests=10, window_seconds=3600
)
# ③ return a 429 with Retry-After when the quota is used up
if blocked:
resp = jsonify({"error": msg})
resp.status_code = 429
resp.headers["Retry-After"] = str(retry_after)
return resp
# Internal service URLs — these use Docker Compose service names, never IPs.
QUICKBITE_URL = "http://quickbite:8000"
SCALERGPT_URL = "http://scalergpt:8000"
DESKBUDDY_URL = "http://deskbuddy-agent:9000"
# Timeout for proxy requests to example services (seconds).
PROXY_TIMEOUT = 30
# ---------------------------------------------------------------------------
# Helper
# ---------------------------------------------------------------------------
def proxy_request(method, url, json_body=None):
"""Forward a request to an internal service and return its JSON response.
Returns a tuple of (response_dict, http_status_code). On connection
errors, returns a helpful error message instead of crashing.
"""
# ① forward the request to the selected internal service
try:
if method == "GET":
resp = http_client.get(url, timeout=PROXY_TIMEOUT)
else:
resp = http_client.post(url, json=json_body, timeout=PROXY_TIMEOUT)
# ② pass through the service JSON and HTTP status code
return resp.json(), resp.status_code
except http_client.ConnectionError:
# ③ turn connection failures into a helpful service-start message
service = url.split("//")[1].split(":")[0]
return {
"error": f"Service '{service}' is not running. "
f"Start it with: docker compose up {service}"
}, 503
except Exception as e:
# ④ return unexpected proxy failures as JSON instead of crashing
return {"error": str(e)}, 500
# ---------------------------------------------------------------------------
# Routes — Static files
# ---------------------------------------------------------------------------
@bp.route("/")
def index():
"""Serve index.html, injecting the correct API base URL for the environment."""
# ① read the static homepage template from the shared src folder
with open(os.path.join(app.static_folder, "index.html"), encoding="utf-8") as f:
html = f.read()
# ② inject the runtime API prefix so browser calls hit this gateway
html = html.replace('data-api-base=""', f'data-api-base="{PATH_PREFIX}"')
# ③ return the rendered HTML with the correct MIME type
return app.response_class(html, mimetype="text/html")
@bp.route("/css/<path:filename>")
def css(filename):
"""Serve stylesheets from the src/css directory."""
return app.send_static_file(os.path.join("css", filename))
@bp.route("/js/<path:filename>")
def js(filename):
"""Serve scripts from the src/js directory."""
return app.send_static_file(os.path.join("js", filename))
@bp.route("/info/<path:filename>")
def info(filename):
"""Serve the "how this demo works" explainer pages from src/info."""
return app.send_static_file(os.path.join("info", filename))
# ---------------------------------------------------------------------------
# Routes — QuickBite ETA (Level 1, keyless)
# ---------------------------------------------------------------------------
@bp.route("/quickbite/predict", methods=["POST"])
def quickbite_predict():
"""Proxy ETA prediction to the QuickBite FastAPI service."""
# ① parse the browser's order JSON
body = request.get_json(force=True)
# ② proxy the order to the QuickBite prediction service
data, status = proxy_request("POST", f"{QUICKBITE_URL}/predict", body)
# ③ return the service JSON and status code unchanged
return jsonify(data), status
@bp.route("/quickbite/status")
def quickbite_status():
"""Check if QuickBite service is running."""
# ① ask QuickBite for its health payload
data, status = proxy_request("GET", f"{QUICKBITE_URL}/")
# ② return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Routes — ScalerGPT (Level 2, needs OPENAI_API_KEY)
# ---------------------------------------------------------------------------
@bp.route("/scalergpt/ask", methods=["POST"])
def scalergpt_ask():
"""Proxy RAG question to the ScalerGPT FastAPI service."""
# ① parse the browser's question JSON
body = request.get_json(force=True)
# ② proxy the question to the ScalerGPT RAG service
data, status = proxy_request("POST", f"{SCALERGPT_URL}/ask", body)
# ③ return the service JSON and status code unchanged
return jsonify(data), status
@bp.route("/scalergpt/status")
def scalergpt_status():
"""Check if ScalerGPT service is running and how many docs are indexed."""
# ① ask ScalerGPT for its health and index summary
data, status = proxy_request("GET", f"{SCALERGPT_URL}/")
# ② return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Routes — DeskBuddy (Level 3, needs OPENAI_API_KEY)
# ---------------------------------------------------------------------------
@bp.route("/deskbuddy/chat", methods=["POST"])
def deskbuddy_chat():
"""Proxy chat message to the DeskBuddy agent service."""
# ① parse the browser's chat JSON
body = request.get_json(force=True)
# ② proxy the message to the DeskBuddy agent loop
data, status = proxy_request("POST", f"{DESKBUDDY_URL}/chat", body)
# ③ return the agent JSON and status code unchanged
return jsonify(data), status
@bp.route("/deskbuddy/status")
def deskbuddy_status():
"""Check if DeskBuddy agent service is running."""
# ① ask DeskBuddy for its health payload
data, status = proxy_request("GET", f"{DESKBUDDY_URL}/")
# ② return the health JSON and status code unchanged
return jsonify(data), status
# ---------------------------------------------------------------------------
# Blueprint registration + server entry point
# ---------------------------------------------------------------------------
app.register_blueprint(bp, url_prefix=PATH_PREFIX)
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)