Key optimizations:
- Short-circuit 'initialize' and 'notifications/initialized' requests
with pre-built static responses, avoiding McpServer/Transport/DB auth
creation for every new MCP session (~2 requests saved per session)
- Add LRU+TTL in-memory cache for API key authentication (5 min TTL)
- Cache read-only DB queries: prompt list, prompt get, search prompts,
get skill, search skills (60-120s TTL depending on operation)
- Pre-serialize GET /api/mcp discovery JSON at module level (once per
cold-start instead of per-request)
- Add Vercel-CDN-Cache-Control and CDN-Cache-Control headers for the
GET endpoint to ensure proper Vercel Edge Network caching
- Add framework-level cache headers in next.config.ts for /api/mcp
- Move rate limiting before server creation so rejected requests
never incur DB auth or server setup costs
- Move body parsing before server creation to enable method inspection
for short-circuiting
Co-authored-by: Fatih Kadir Akın <fka@fka.dev>
Learn prompt engineering with our free, interactive guide — 25+ chapters covering everything from basics to advanced techniques like chain-of-thought reasoning, few-shot learning, and AI agents.