> From https://dialect-engineering.ai/sections/11-degrees-of-freedom, the companion site to the paper "SaaS architecture when code is cheap". Draft, September 2026.

# SaaS architecture when code is cheap

## 11. Degrees of freedom and what they cost

A narrower target performs better for a reason that also predicts where the cost goes.

| Layer | Degrees of freedom |
|---|---|
| Natural language | Unbounded. Meaning is completed by the reader |
| General-purpose code | Near unbounded |
| A domain grammar | Bounded by the grammar |
| A schema or API surface | Very narrow and fixed |

A model translates natural language into general-purpose code comparatively well because the two have similar freedom, so the crossing is short. Asking a model to go from natural language to a narrow schema in one step is a long crossing, and the model performs the narrowing invisibly: it picks a reading, the software executes it, and a misread request returns something plausible with no record of which reading was chosen. A grammar splits one long crossing into two shorter ones and makes the intermediate result inspectable.

The same argument has a cost consequence. Getting from prose instructions to a valid structure takes the model work, and that work is tokens. The narrower and more explicit the target, the less of it there is. In one of our own applications, moving a configuration from a schema-plus-instructions arrangement to a grammar reduced token consumption substantially. The reduction was observed in use and has not been measured under controlled conditions, so no figure is given.

A related measurement concerns context cost. A reproducible benchmark across MCP servers found a 25x spread in the token cost of tool definitions, with one official server consuming 17,161 tokens for 24 tools and 97% of that cost coming from input schemas rather than descriptions ([mcp-token-benchmark](https://github.com/zhang-liz/mcp-token-benchmark)). Inference prices, though, are falling fast: Epoch AI finds the price for a fixed capability level falling by a median of 50x per year ([Epoch AI](https://epoch.ai/data-insights/llm-inference-price-trends)). So the durable part of this argument is review load rather than token spend.

Both matter for sequencing rather than for the pitch. At pilot scale, neither token cost nor review load is noticed. Both become visible at the point where a working pilot is asked to cover the whole product.

---

---

Previous: [10. Configuration schema and grammar](https://dialect-engineering.ai/sections/10-schema-and-grammar.md) · Next: [12. Verification moves from reading to running](https://dialect-engineering.ai/sections/12-verification.md)
