Two a week for six weeks, each carrying one capability the market is hiring for. None of this is finished — that is what makes it a roadmap rather than a portfolio, and it is why it lives here instead of on the front page.
The constraint is the point.
Twelve standalone demos prove twelve times that a model can be called from an application. One platform that twelve systems run on proves something harder and more relevant: that a model gateway, observability, evals, spend caps, and local-plus-frontier inference can be built once and reused — which is the actual shape of the problem inside an enterprise.
So every system below sits on the same core. A model swap is a gateway change, not twelve changes. A cost regression is visible in one trace store. An eval gate applies to all of them or to none. If the platform is wrong, twelve systems make that obvious quickly — which is the cheapest way to find out.
Ordered by week. Status is honest: two started, ten not.