Performance & Cost Transparency

A clear view of what a full production build consumes — and the optimizations that keep it efficient. These are the real numbers from the HCP 360 Oncology Context engagement.

Engagement Cost Summary

A complete requirements-to-dbt build — a 24-table Snowflake pipeline designed, coded, and validated end to end

40.95M
Total tokens, across 4 pipeline stages
18.5
Platform credits consumed
$55.50
Total cost for the full run
24
Tables built — intermediates, dimensions, spine & consumption

Token Usage by Pipeline Stage

Four stages, one engagement — with the entities each stage processed and the inputs it worked from

Stage Tokens Entities processed Key inputs
Requirements Agent 9.87M
9,870,878
24 pipeline tables scoped — 13 intermediates, dimension & reference tables, the HCP_UNIVERSE spine, and the HCP_360 consumption table 22 extracted files — 6 working-session transcripts (~65 pages), 10 legacy SQL scripts, an 825-line technical metadata spreadsheet, and a 392-line source entity-relationship model
STTM Generation + Validation 9.21M
9,208,987
24 target entities across 6 dependency tiers BRD, FRD and data requirements from the Requirements Agent, 44 parsed data-model CSVs, and the entity-relationship graph
Code Generation + Validation + Dashboard Scoring 16.95M
16,952,147
24 tables — 17 via SQL DML reuse, 7 generated from STTM skeletons The 400-line reviewed STTM, topological lineage order, the 6-tier pipeline design, and per-entity metadata — producing 24 DDL and 24 DML scripts
dbt Pipeline Conversion 5.01M
5,011,843
19 dbt models — 6 staging, 12 intermediate, 1 marts The 24 DDL and 24 DML scripts, reviewed STTM, and project configuration — 26/26 models passed structural validation
Total 40.95M
41,043,855
18.5 credits · $55.50 total

What the Engagement Delivered

A complete, documented, and validated data product — not just code

87
KPIs documented across 14 metric families
66
Functional requirements, fully traced
34
User stories across 7 personas
31
Data-quality rules defined
15
Business requirements in the BRD
11
Key business questions across 5 domains
24 + 24
DDL and DML scripts, one pair per table
6 + 6
Markdown deliverables + branded DOCX documents

How Token Usage Is Optimized

A generic LLM pipeline re-sends and re-generates everything, every run. OneData is built around one principle: a token is only spent where a model call genuinely adds value

🔑

SHA-256 change detection

Every input document, mapping row, and generated script is fingerprinted into a version manifest. On a re-run, an unchanged hash means the stored output is reused — the model is never called for work that's already done.

⚙️

Deterministic tools, not model calls

Steps with one correct answer — document parsing, DDL generation, Excel conversion, dbt scaffolding, DOCX branding, validation scoring — run as plain scripts through a unified pipeline CLI. LLM tokens are reserved for genuinely ambiguous work like transformation logic.

♻️

Legacy SQL reuse

Where a trusted production script already exists, it's ingested and adapted rather than regenerated from scratch. In the HCP 360 run, 17 of the 24 tables were built through SQL DML reuse — only 7 needed fresh generation from STTM skeletons.

🔗

Lineage-aware, per-entity context

Tables are ordered into a topological lineage, and each one gets its own assembled context bundle — the mapping rows, joins, and metadata it actually needs. The model sees the relevant slice, never the whole corpus, on every call.

🎯

Incremental updates, not re-runs

When inputs change, code is hand-edited, or a change request lands, a dedicated agent classifies real vs. cosmetic changes, computes the lineage-aware blast radius, and surgically regenerates only the affected entities — never the full pipeline.

📦

Versioned snapshots

Every completed phase is archived as an immutable version with its manifest. Later runs diff against the last snapshot and start from the current state instead of starting cold — re-work is measured, not assumed.

🔁

How routing works. When an entity reaches a stage, its SHA-256 fingerprint is compared against the version manifest from the last archived run. A match means the stored output is reused with no model call; a mismatch routes only that entity — with only its context bundle — to the model. Everything else stays untouched.

OneData vs. a generic LLM pipeline

Same job, same entities — the difference is what happens before a token is ever spent

OneData Generic LLM pipeline
Change detection ✅ SHA-256 fingerprint per input, mapping row, and script — only real changes reach the model ❌ Full re-send and re-generation on every run
Deterministic work ✅ Scripted through a unified CLI — parsing, DDL, Excel, dbt scaffolding, validation scoring ❌ Routed through the model regardless of ambiguity
Existing code ✅ Trusted legacy SQL is ingested and reused — 17 of 24 HCP 360 tables came through DML reuse ❌ Ignored — everything regenerated from scratch
Handling change ✅ Blast-radius analysis regenerates only the affected entities ❌ Any change means a full pipeline re-run
Correctness check ✅ Completeness score + scoring dashboard, built into the pipeline ❌ Manual QA after the fact
Cost visibility ✅ Per-stage token, credit, and dollar ledger ❌ Opaque aggregate bill
⚖️

Optimized for tokens, not for correctness cut corners. Every generated artifact is scored before it's called done — a completeness score checks every required field, table, and mapping; a Streamlit dashboard shows how generated code scores against validation rules per entity; and failed checks route back into generation automatically instead of surfacing as silent gaps downstream. See how validation works →

Bring OneData to Your Engagement

Every engagement comes with full per-stage cost reporting — tokens, credits, and dollars. Talk to us about what a run would look like on your data.

Get in Touch