← Shwetal Mehta
Roadmap

Twelve systems, one platform

Two a week for six weeks, each carrying one capability the market is hiring for. None of this is finished — that is what makes it a roadmap rather than a portfolio, and it is why it lives here instead of on the front page.

Why one platform and not twelve demos

The constraint is the point.

Twelve standalone demos prove twelve times that a model can be called from an application. One platform that twelve systems run on proves something harder and more relevant: that a model gateway, observability, evals, spend caps, and local-plus-frontier inference can be built once and reused — which is the actual shape of the problem inside an enterprise.

So every system below sits on the same core. A model swap is a gateway change, not twelve changes. A cost regression is visible in one trace store. An eval gate applies to all of them or to none. If the platform is wrong, twelve systems make that obvious quickly — which is the cheapest way to find out.

The board

Ordered by week. Status is honest: two started, ten not.

WK 1Talent Scout — agentic job search: bot surface, scheduled runs, human approval gate, calibrated fit scoresIn build WK 1Platform Core — model gateway across Anthropic, OpenAI, Google, xAI and local models; traces and spend capsIn build WK 2Upgrade Gate — evals-as-CI; cross-vendor A/B on quality, cost and latency; go/no-go on every model swapQueued WK 2Local Inference — open-weight models on-prem behind the gateway; local triage, frontier judgmentQueued WK 3Governance Copilot — agentic retrieval with citations over NIST AI RMF, the EU AI Act and ISO 42001; multimodal ingestionQueued WK 3ITSM MCP Server — a real ticketing system exposed to agents with identity, scoped permissions and an audit logQueued WK 4NOC Triage Agent — real-time triage over a streaming alert feed; reads dashboards and topology; human-in-the-loop on consequential actionsQueued WK 4Voice Service Desk — call in, describe the problem, get a ticket and a status; realtime voice with human handoffQueued WK 5Red Team & Harden — OWASP Agentic Top 10 and MCP Top 10 against the platform; published fix reportQueued WK 5Three Ways to Make a Decision — prompted frontier model vs fine-tuned open model vs decision model, scored on accuracy, calibration, latency and costQueued WK 6One Workload, Two Clouds — the same system on Bedrock AgentCore and Azure AI FoundryQueued WK 6Research Factory — overnight multi-agent research with cross-family rebuttal; morning brief as text and audioQueued

Elsewhere