Once Cortex Code is set up, running OneData is a guided, phase-by-phase flow. You launch the CLI from the project root, and the agents move through each phase — pausing for your approval at every gate.
OneData has a reusable accelerator that travels across every client, and a per-project workspace that holds one engagement's inputs and outputs. You own the inputs; OneData creates everything else.
.cortex/ ├─ agents/ # the specialist agents (identity + allowed tools) │ ├─ requirement-generation-agent.md │ ├─ design-agent.md │ ├─ sttm-generation-agent.md │ ├─ sttm-validation-agent.md │ ├─ codegen-agent.md │ ├─ code-validation-agent.md │ ├─ dbt-pipeline-agent.md │ ├─ incremental-update-agent.md │ └─ ... (validators + dbt coworkers) └─ skills/ # workflow instructions + governing rules ├─ requirement-generation-workflow/ ├─ design-workflow/ ├─ sttm-generation-workflow/ ├─ codegen-workflow/ ├─ dbt-pipeline-workflow/ └─ rules/ # security, naming, Snowflake, STTM, dbt, DQ
<Project_Name>/ ├─ config/project_config.json # database, schema, warehouse, role ├─ scripts/pipeline_cli.py # single entry point for every tool ├─ input/ # ◀ YOU provide these │ ├─ raw/ # transcripts, docs, SQL, contracts │ ├─ brd/ # business rules (.docx) │ ├─ models/ # technical metadata (.xlsx) │ ├─ er/ # entity relationships │ └─ change_requests/ # drop CR memos here for updates └─ output/ # ▶ OneData creates these ├─ requirements/ # BRD, KPIs, KBQs ├─ design/ # ARD/RRD, pipeline design, models ├─ sttm/ # mappings (CSV + Excel) ├─ generated/ # DDL, DML, scores ├─ versions/ # immutable snapshots └─ logs/
You only touch input/. Drop your requirements and metadata in, and the agents create and update everything under output/ — phase by phase, with approval gates and validation between each.
OneData detects what you've provided and starts at the right phase — you don't always begin from scratch
| Scenario | You provide | Pipeline starts at |
|---|---|---|
| 1 | Raw evidence only (transcripts, docs, SQL) | Requirements Generation Agent → Design → … |
| 2 | Raw evidence + draft ARD/RRD designs | Design validates your drafts, then continues |
| 3 | Raw evidence + designs + pipeline plan | Design fills the table-level detail, then continues |
| 4 | Full inputs (BRD + models + ER) | Straight to STTM Generation Agent → Code → dbt |
Scenario detection is automatic — a quick detect-scenario check reads your input/ folder and picks the right entry point.
Each phase does its work, asks for what it needs, and produces a reviewable artifact before the next begins
.cortex package, creates the output folders, and reads your inputs.Everything runs through one entry point, pipeline_cli.py. A typical run looks like this — the agents drive it, so you mostly approve gates.
# Detect what you provided and set up the workspace python scripts/pipeline_cli.py extract-inputs python scripts/pipeline_cli.py detect-scenario # Requirements & design run as agents, with approval gates … # then, from mappings onward: python scripts/pipeline_cli.py init python scripts/pipeline_cli.py discover python scripts/pipeline_cli.py validate-sttm # independent check python scripts/pipeline_cli.py excel # review-ready workbook # After sign-off: generate, score, package python scripts/pipeline_cli.py generate-ddl <ENTITY> python scripts/pipeline_cli.py score <ENTITY> python scripts/pipeline_cli.py dbt # Version the completed run for safe incremental updates python scripts/pipeline_cli.py snapshot python scripts/pipeline_cli.py archive