A clear view of what a full production build consumes — and the optimizations that keep it efficient. These are the real numbers from the HCP 360 Oncology Context engagement.
A complete requirements-to-dbt build — a 24-table Snowflake pipeline designed, coded, and validated end to end
Four stages, one engagement — with the entities each stage processed and the inputs it worked from
| Stage | Tokens | Entities processed | Key inputs |
|---|---|---|---|
| Requirements Agent | 9.87M 9,870,878 |
24 pipeline tables scoped — 13 intermediates, dimension & reference tables, the HCP_UNIVERSE spine, and the HCP_360 consumption table | 22 extracted files — 6 working-session transcripts (~65 pages), 10 legacy SQL scripts, an 825-line technical metadata spreadsheet, and a 392-line source entity-relationship model |
| STTM Generation + Validation | 9.21M 9,208,987 |
24 target entities across 6 dependency tiers | BRD, FRD and data requirements from the Requirements Agent, 44 parsed data-model CSVs, and the entity-relationship graph |
| Code Generation + Validation + Dashboard Scoring | 16.95M 16,952,147 |
24 tables — 17 via SQL DML reuse, 7 generated from STTM skeletons | The 400-line reviewed STTM, topological lineage order, the 6-tier pipeline design, and per-entity metadata — producing 24 DDL and 24 DML scripts |
| dbt Pipeline Conversion | 5.01M 5,011,843 |
19 dbt models — 6 staging, 12 intermediate, 1 marts | The 24 DDL and 24 DML scripts, reviewed STTM, and project configuration — 26/26 models passed structural validation |
| Total | 40.95M 41,043,855 |
18.5 credits · $55.50 total | |
A complete, documented, and validated data product — not just code
A generic LLM pipeline re-sends and re-generates everything, every run. OneData is built around one principle: a token is only spent where a model call genuinely adds value
Every input document, mapping row, and generated script is fingerprinted into a version manifest. On a re-run, an unchanged hash means the stored output is reused — the model is never called for work that's already done.
Steps with one correct answer — document parsing, DDL generation, Excel conversion, dbt scaffolding, DOCX branding, validation scoring — run as plain scripts through a unified pipeline CLI. LLM tokens are reserved for genuinely ambiguous work like transformation logic.
Where a trusted production script already exists, it's ingested and adapted rather than regenerated from scratch. In the HCP 360 run, 17 of the 24 tables were built through SQL DML reuse — only 7 needed fresh generation from STTM skeletons.
Tables are ordered into a topological lineage, and each one gets its own assembled context bundle — the mapping rows, joins, and metadata it actually needs. The model sees the relevant slice, never the whole corpus, on every call.
When inputs change, code is hand-edited, or a change request lands, a dedicated agent classifies real vs. cosmetic changes, computes the lineage-aware blast radius, and surgically regenerates only the affected entities — never the full pipeline.
Every completed phase is archived as an immutable version with its manifest. Later runs diff against the last snapshot and start from the current state instead of starting cold — re-work is measured, not assumed.
How routing works. When an entity reaches a stage, its SHA-256 fingerprint is compared against the version manifest from the last archived run. A match means the stored output is reused with no model call; a mismatch routes only that entity — with only its context bundle — to the model. Everything else stays untouched.
Same job, same entities — the difference is what happens before a token is ever spent
| OneData | Generic LLM pipeline | |
|---|---|---|
| Change detection | ✅ SHA-256 fingerprint per input, mapping row, and script — only real changes reach the model | ❌ Full re-send and re-generation on every run |
| Deterministic work | ✅ Scripted through a unified CLI — parsing, DDL, Excel, dbt scaffolding, validation scoring | ❌ Routed through the model regardless of ambiguity |
| Existing code | ✅ Trusted legacy SQL is ingested and reused — 17 of 24 HCP 360 tables came through DML reuse | ❌ Ignored — everything regenerated from scratch |
| Handling change | ✅ Blast-radius analysis regenerates only the affected entities | ❌ Any change means a full pipeline re-run |
| Correctness check | ✅ Completeness score + scoring dashboard, built into the pipeline | ❌ Manual QA after the fact |
| Cost visibility | ✅ Per-stage token, credit, and dollar ledger | ❌ Opaque aggregate bill |
Optimized for tokens, not for correctness cut corners. Every generated artifact is scored before it's called done — a completeness score checks every required field, table, and mapping; a Streamlit dashboard shows how generated code scores against validation rules per entity; and failed checks route back into generation automatically instead of surfacing as silent gaps downstream. See how validation works →
Every engagement comes with full per-stage cost reporting — tokens, credits, and dollars. Talk to us about what a run would look like on your data.
Get in Touch