Featured

Deploy OpenClaw in 60 seconds — 20% off logoDeploy OpenClaw in 60 seconds — 20% off

Launch OpenClaw on Hostinger in about 60 seconds and keep your agent live 24/7. Our referral link gives you 20% off, no coupon code needed.

Launch on Hostinger
Run your Hermes agent on Hostinger, fully managed logoRun your Hermes agent on Hostinger, fully managed

Launch Hermes on Hostinger in one click, fully managed, no VPS knowledge needed. Use code ZACAARON10 for 10% off.

Launch on Hostinger
Crawl and scrape any site into clean data, 10% off logoCrawl and scrape any site into clean data, 10% off

Firecrawl crawls and scrapes any site into clean markdown for your agent. Get 1,000 free credits, and new users get 10% off their first purchase.

Try Firecrawl free
Your own AI agent, running 24/7 with QwikClaw logoYour own AI agent, running 24/7 with QwikClaw

QwikClaw sets up and runs an always-on OpenClaw agent for you. One click, no config files, no server setup.

Deploy now
One API to scrape, enrich, and extract the internet. logoOne API to scrape, enrich, and extract the internet.

Context.dev gives your agents a single API to scrape, enrich, and extract live web data — no proxies, no parsers, no maintenance.

Start building free
SetupClaw: done-for-you OpenClaw for founders & exec teams logoSetupClaw: done-for-you OpenClaw for founders & exec teams

White-glove OpenClaw for founders and exec teams (4–50+ employees): we install, harden, integrate your tools, and maintain it — secured from day one.

Get it set up for you
SEO data APIs for your agent, $1 free credit logoSEO data APIs for your agent, $1 free credit

DataForSEO gives your agent live access to SERP results, keyword data, backlinks, and on-page SEO data through one API. New accounts get a $1 credit, good for up to 20,000 keyword or backlink lookups.

Try DataForSEO free
Reach 47,000+ AI builders

A flat monthly placement in front of developers actively installing AI tools. No lock-in, cancel anytime.

Advertise here

Works with

Claude CodeClaude DesktopCursorVS CodeClineCodex CLIOpenClaw+ any MCP client

Install to Claude Code

This server doesn't publish a one-line install command. Follow the setup in the source repository.

Summary

A custom MCP server providing tools for date/time, calculations, mock weather, and note management, enabling AI agents to perform these tasks via natural language.

README.md

MCP Learning Project

A hands-on project to learn Model Context Protocol (MCP) by building a custom MCP server, an AI agent, and a full-stack web application with semantic document search, model routing, prompt evaluation, and a live AI cost observability dashboard.

---

What is MCP?

Model Context Protocol (MCP) is an open standard that lets AI models (like Claude) call external tools and services in a structured, language-agnostic way. Think of it like USB — any tool built to the MCP standard works with any MCP-compatible AI.

---

Project Structure

MCP Project/
├── src/
│   ├── backend/
│   │   ├── api.py                  — FastAPI web server (primary entry point)
│   │   ├── agent.py                — CLI agent (original learning version)
│   │   ├── mcp_server.py            — MCP server: 8 tools, 2 resource kinds, 1 prompt
│   │   ├── text_editor_tool.py      — Client-side text editor tool, locked to knowledge_base/project_notes.md
│   │   ├── database.py              — SQLite layer (notes, sessions, usage_logs, credit_config)
│   │   └── rag.py                   — ChromaDB semantic search
│   └── frontend/
│       └── chat.html            — Browser chat UI (SSE streaming)
├── scripts/
│   ├── convert_pdfs.py          — Tesseract OCR for scanned PDFs
│   ├── inspect_db.py            — Utility to view SQLite contents
│   └── tool_use_demo.py         — Tool Use Fundamentals demo (WARNING: consumes API credits)
├── knowledge_base/          — Drop your documents here (RAG source content)
├── data/                    — data.db (SQLite) + chroma_db/ (ChromaDB), both gitignored
├── evals/
│   ├── dataset.json         — 12 test cases for tool selection + model routing
│   └── run_evals.py         — Eval runner (WARNING: consumes API credits)
├── docs/                    — Project documentation (see table below)
├── .mcp.json                — Project-scoped MCP servers (Playwright, for UI testing)
├── README.md
├── CLAUDE.md                — Instructions for Claude Code
└── requirements.txt

---

Architecture

Browser (http://localhost:8000)
  │
  │ HTTP / Server-Sent Events
  ▼
api.py (FastAPI)
  │
  ├──► Claude Sonnet 4.6 / Haiku 4.5 (Anthropic API — routed by query complexity)
  │         │ tool calls
  │         ▼
  ├──► mcp_server.py (8 MCP Tools, 2 resource kinds, 1 prompt)
  │         ├──► database.py  → SQLite (notes, sessions, usage_logs, credit_config)
  │         ├──► rag.py       → ChromaDB (semantic document search)
  │         └──► knowledge_base/ → your documents (txt, md, PDF)
  │
  ├──► text_editor_tool.py (client-side tool, in-process — locked to knowledge_base/project_notes.md)
  │
  └──► web_search (server-side tool — runs on Anthropic's infrastructure, no local code)

Three processes run together: the browser, api.py, and mcp_server.py (spawned as a subprocess and kept alive for the life of the app). See docs/ARCHITECTURE.md for the full request lifecycle.

---

All 10 Tools

Three different execution models, one tools list:

| Tool | Execution | Description | |---|---|---| | get_current_datetime | MCP | Current date and time | | calculate | MCP | Safe math expression evaluator | | get_weather | MCP | Mock weather data by city | | manage_notes | MCP | Persistent CRUD notes (SQLite) | | list_docs | MCP | Lists files in knowledge_base/ folder | | read_doc | MCP | Reads full content of a document | | index_docs | MCP | Indexes docs into ChromaDB for semantic search | | search_docs | MCP | Semantic search — finds relevant chunks for any query | | web_search | Server-side (Anthropic) | Live web search for anything time-sensitive or beyond training data. $10/1,000 searches + token cost. | | str_replace_based_edit_tool | Client-side (local) | Lets Claude view/edit exactly one file — knowledge_base/project_notes.md — nothing else |

---

MCP Resources & Prompts

mcp_server.py uses all three MCP primitives, not just tools:

  • Resources — read-only, URI-addressable data: knowledgebase://files (the file listing, as a resource instead of a tool call) and note://<title>, one per saved note.
  • Prompts — reusable request templates: summarize_document takes a filename and returns a pre-built request that drives the existing read_doc/search_docs tools.

Reachable through the running app, not just a standalone test script — GET /resources, GET /resources/content, GET /prompts, POST /prompts/{name}. See CLAUDE.md § MCP Resources & Prompts for the full design, including a real gotcha (URI schemes can't contain underscores — RFC 3986) caught via testing, and a real wiring gap (found via /code-review, fixed) where these worked in isolated tests while being unreachable in production.

---

Image + PDF Attachments

Attach an image or PDF to a chat message (📎 button in the chat UI) and Claude reads it directly — native vision/document understanding, no OCR pipeline needed. This is for ad-hoc content brought into a conversation (a screenshot to debug, a one-off PDF to summarize), separate from the curated, pre-indexed knowledge_base/ corpus used by search_docs.

  • Ephemeral — the file is sent to Claude for that turn only; conversation history persists as

plain text, never the binary, so later turns don't resend it.

  • Citations — PDF attachments get page citations enabled; when Claude cites specific content,

the response includes an inline (p.N) page reference.

  • No new cost tracking needed — image/PDF tokens bill as ordinary input tokens, already

captured the same way every other request's usage is.

  • Limits served from the backendGET /attachment-limits is the single source of truth

for allowed file types and size caps; the frontend fetches it on load instead of hardcoding a second copy that could drift out of sync.

See CLAUDE.md § Image + PDF Attachments for the full design (validation, size caps, the ephemeral-history mechanism).

---

Setup

Prerequisites

  • Python 3.10+
  • An Anthropic API key (console.anthropic.com)
  • Tesseract OCR (for scanned PDFs): github.com/UB-Mannheim/tesseract/wiki

Install dependencies

pip install -r requirements.txt

Set your API key

Create a .env file in the project root: `` ANTHROPIC_API_KEY=sk-ant-... ``

Windows tip: save .env as plain UTF-8 (no BOM). Notepad and some PowerShell commands default to "UTF-8 with BOM," which silently breaks python-dotenv and produces a "Could not resolve authentication method" error even though the key is correct.

Environments (optional)

No setup needed to just run the app — the default development environment behaves exactly as above. To run under a different environment, set ENVIRONMENT and create a matching .env.<environment> file (e.g. .env.production); it gets its own SQLite database file too (data.<environment>.db), so it can never share data with local dev. See CLAUDE.md for details.

Run the web app

python -m uvicorn api:app --reload --port 8000 --app-dir src/backend

Open http://localhost:8000 for the chat UI. Cost/usage tracking and alerting are handled by SpendGaugeAI — see below.

Or run the CLI agent

python src/backend/agent.py

---

AI Cost Dashboard

This project's own local /usage dashboard (token/cost/session tracking, credit balance, Discord alerts) was removed 2026-07-19 in favor of SpendGaugeAI — the standalone, general-purpose extraction of this exact dashboard, with the same feature set (4-way token breakdown, cost by model/project/tool, credit balance tracker with burn rate/runway, and the same Discord alert types) plus multi-app support. This project reports to it via SPENDGAUGEAI_URL/ SPENDGAUGEAI_API_KEY in .env — see § Logging, Errors & Tracing below for setup. Tracks Anthropic API usage only — not your Claude Pro subscription (a separate, flat-fee product).

---

Logging, Errors & Tracing

Visit /logs for two tabs:

  • Logs — errors/warnings/info parsed from data/app.log, click-to-filter summary cards, expandable full tracebacks
  • Conversations — the actual content of recent chats (not just cost/token metadata), reading directly from the existing session history

Structured logging (Python's logging module) writes to both the console and a rotating data/app.log (14-day retention). Request latency is tracked at every /chat//stream exit point, success and failure alike.

Optional: Langfuse tracing. Set LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY in .env (free tier at cloud.langfuse.com) to get full end-to-end tracing of every Claude API call, viewable in Langfuse's own dashboard. Fully optional — every call site no-ops cleanly if unset, same pattern as the Discord webhook. See CLAUDE.md § Logging & Tracing for the full design.

SpendGaugeAI reporting. Set SPENDGAUGEAI_URL and SPENDGAUGEAI_API_KEY in .env and pip install spendgaugeai to report every request's usage to a running SpendGaugeAI instance — this project's own extraction, now the only place usage/cost/credit/alerts are tracked (this project's own local /usage dashboard was removed 2026-07-19; see § AI Cost Dashboard above). See CLAUDE.md § Logging & Tracing for the full design.

SpendGaugeAI hard budget enforcement (opt-in). Set SPENDGAUGEAI_ENFORCE=1 in .env (on top of the two vars above) to actually block /chat and /stream with a 402 once the budget you've set on the SpendGaugeAI dashboard is exhausted, instead of only reporting after the fact. A separate flag from the reporting vars on purpose — blocking a real chat response is a user-visible behavior change, so it needs its own explicit opt-in. Fails open (never blocks) if SpendGaugeAI is unreachable or the check times out. Off by default; unset means today's existing report-only behavior is unchanged.

---

Model Routing & Prompt Caching

Not every message needs the same model. _pick_model() routes short/simple queries to Haiku (10–20× cheaper) and long or document-related queries to Sonnet. The system prompt is marked cache_control: ephemeral, saving ~90% of its token cost after the first call in a 5-minute window. See LEARNING_JOURNEY.md Phase 8–9 for the full breakdown.

---

Eval Pipeline

12 test cases verify Claude follows system prompt rules — correct tool selection and correct model routing — scored automatically.

# Start the app first, then in a second terminal:
python evals/run_evals.py

Cost warning: each eval case makes a real Claude API call. 12 cases = 12 API calls.

Currently passing: 12/12 (100%). Run after every system prompt or routing change.

---

UI Testing with Playwright MCP

A Playwright MCP server (Microsoft's official @playwright/mcp) is configured at project scope in .mcp.json, letting Claude Code drive chat.html and usage.html in a real browser — navigate, click, type, screenshot — instead of only reading source code to guess whether a UI change works.

claude mcp add playwright --scope project -- npx -y @playwright/mcp

Restart Claude Code (or reconnect MCP servers) after adding it, then verify with /mcp.

No secrets are needed for this server — safe to commit .mcp.json and share across a team. Screenshots and page snapshots it produces land in .playwright-mcp/, which is gitignored since they can capture session IDs and cost data from local testing.

This is more than a nice-to-have: a live end-to-end test of the chat UI is what caught the .env UTF-8 BOM bug documented above — a bug that pure code review would have missed entirely, since the API key was correct and the failure only appeared once a real request was made.

---

How to Add a New Tool

MCP tool (needs local execution logic — most tools):

Step 1 — Declare the tool in list_tools() inside mcp_server.py: ``python types.Tool( name="my_tool", description="What it does and WHEN Claude should use it.", inputSchema={"type": "object", "properties": {"param": {"type": "string"}}, "required": ["param"]}, ), ``

Step 2 — Handle it in call_tool() inside mcp_server.py: ``python if name == "my_tool": result = do_something(arguments["param"]) return [types.TextContent(type="text", text=result)] ``

Restart the server — Claude discovers the new tool automatically.

Anthropic-native tool (server-side like web_search, or client-side like the text editor): these bypass MCP entirely and are declared directly in app.state.tools inside api.py's lifespan — see CLAUDE.md § Adding a New Tool for the server-side vs. client-side pattern and the allowed_callers/model-capability gotcha that broke Haiku-routed requests during this project's web_search rollout.

---

How to Add Documents

  1. Drop .txt, .md, or .pdf files into the knowledge_base/ folder
  2. For scanned PDFs: run python scripts/convert_pdfs.py first
  3. Restart the server (auto-indexes on startup) or say "Re-index my documents" in chat

---

RAG — How Semantic Search Works

Indexing (once):
  knowledge_base/*.txt → split into ~500 char chunks → embed with all-MiniLM-L6-v2 → store in ChromaDB

Querying (every question):
  question → embed → ChromaDB similarity search → top 4 relevant chunks → Claude

This handles documents of any size — only the relevant parts are sent to Claude.

---

Key Concepts

| Concept | File | Purpose | |---|---|---| | @app.list_tools() | mcp_server.py | Declares tools to any MCP client | | @app.call_tool() | mcp_server.py | Executes tools and returns results | | lifespan | api.py | Keeps MCP server alive across all HTTP requests | | StreamingResponse | api.py | SSE streaming to the browser | | _pick_model() | api.py | Routes each message to Haiku or Sonnet | | _spendgauge_report() | api.py | Reports tokens, cost, tools called per request to SpendGaugeAI | | init_db() | database.py | Creates SQLite tables on startup | | index_all() | rag.py | Chunks + embeds all docs into ChromaDB | | search() | rag.py | Semantic similarity search | | async_mcp_tool() | agent.py / api.py | Bridges MCP tools to Anthropic SDK | | BetaAsyncBuiltinFunctionTool | text_editor_tool.py | SDK interface for client-side Anthropic tools — to_dict() + call() |

---

Dependencies

| Package | Purpose | |---|---| | anthropic[mcp] | Anthropic SDK + MCP integration | | mcp | MCP protocol implementation | | fastapi | Web framework | | uvicorn[standard] | ASGI web server | | pypdf | Text-based PDF extraction | | pymupdf | PDF → image rendering for OCR | | pytesseract | Tesseract OCR wrapper | | chromadb | Vector database | | sentence-transformers | Local embedding model | | python-dotenv | Loads .env into environment variables | | httpx | HTTP client (Anthropic SDK + eval runner) |

---

Windows SSL Note

Two SSL patches are applied for Windows machines with corporate certificate chains or network monitoring drivers: rag.py defaults httpx client verify=False for the embedding model download, and api.py's lifespan clears SSLKEYLOGFILE and passes an explicit httpx.AsyncClient (verify=False) to AsyncAnthropic(). See CLAUDE.md for details.

---

Documentation

README.md and CLAUDE.md stay at the project root; the rest live in docs/:

| File | Purpose | |---|---| | CLAUDE.md (root) | Instructions for Claude Code — commands, architecture, standards | | docs/ARCHITECTURE.md | System design in plain English | | docs/LEARNING_JOURNEY.md | Phase-by-phase build record | | docs/LEARNING_PLAN.md | Roadmap to expert AI engineer | | docs/INSIGHTS.md | Key lessons and principles | | docs/TUTORIAL.md | Beginner teaching guide with exercises | | docs/GIT_COMMANDS.md | All Git commands used, with explanations | | docs/AI_ENGINEERING_PORTFOLIO.md | Skills portfolio for hiring managers |

---

GitHub

github.com/vijayanan6/mcp-project

---

What's Next

See LEARNING_PLAN.md for the full roadmap. Near-term:

  • pytest — unit + integration tests for MCP tools and API routes
  • Docker + GCP Cloud Run deployment
  • PostgreSQL (replacing SQLite) + pgvector (replacing ChromaDB)
  • React frontend with authentication (JWT)
  • Multi-model support — Gemini, OpenAI, and free local models via Ollama/Groq
  • Multi-agent systems and a second project in a different domain

See related servers & alternatives →

Related MCP servers

Browse all →

Related guides

Hand-picked reading to help you choose and use Productivity servers.