AI Router
An exploration into AI engineering, building a smart router that directs requests to the most appropriate LLM based on cost, speed, and capability.

01. The Problem
Calling the most powerful LLM for every request is expensive and slow. Simple tasks like summarisation don't need GPT-4-level capability, but developers lack a simple way to route intelligently without building custom logic for every project.
02. The Solution
A drop-in API router that accepts a standard chat-completion request, scores the prompt on complexity and required capability, selects the optimal model, forwards the request, and returns the response — all transparently to the client.
03. Architecture & Tech Decisions
A Node.js/Express middleware layer that classifies incoming prompts by complexity and intent, then routes them to the appropriate model (GPT-4o for complex reasoning, Claude Haiku for speed, GPT-3.5 for simple tasks) using a scoring system.
04. Testing Strategy
Tested against a benchmark suite of 50 prompts spanning simple, medium, and complex categories. Measured accuracy of model selection, latency, and cost-per-request compared to always using GPT-4o.
05. Deployment
Deployed as a standalone Express server. Designed to be self-hosted or run as a sidecar service alongside any application that consumes LLM APIs.
06. Lessons Learned
Simplicity beats cleverness in routing logic. An initial approach using a secondary LLM to classify prompts added cost and latency that negated the savings. The heuristic approach was 80% as accurate at 0.1% of the cost.