▶️

Running the Pipeline

Once Cortex Code is set up, running OneData is a guided, phase-by-phase flow. You launch the CLI from the project root, and the agents move through each phase — pausing for your approval at every gate.

Two folders, one idea

OneData has a reusable accelerator that travels across every client, and a per-project workspace that holds one engagement's inputs and outputs. You own the inputs; OneData creates everything else.

.cortex/ — the reusable accelerator
.cortex/
├─ agents/                        # the specialist agents (identity + allowed tools)
│  ├─ requirement-generation-agent.md
│  ├─ design-agent.md
│  ├─ sttm-generation-agent.md
│  ├─ sttm-validation-agent.md
│  ├─ codegen-agent.md
│  ├─ code-validation-agent.md
│  ├─ dbt-pipeline-agent.md
│  ├─ incremental-update-agent.md
│  └─ ... (validators + dbt coworkers)
└─ skills/                        # workflow instructions + governing rules
   ├─ requirement-generation-workflow/
   ├─ design-workflow/
   ├─ sttm-generation-workflow/
   ├─ codegen-workflow/
   ├─ dbt-pipeline-workflow/
   └─ rules/                      # security, naming, Snowflake, STTM, dbt, DQ
<Project>/ — one engagement
<Project_Name>/
├─ config/project_config.json     # database, schema, warehouse, role
├─ scripts/pipeline_cli.py        # single entry point for every tool
├─ input/                         # ◀ YOU provide these
│  ├─ raw/                        # transcripts, docs, SQL, contracts
│  ├─ brd/                        # business rules (.docx)
│  ├─ models/                     # technical metadata (.xlsx)
│  ├─ er/                         # entity relationships
│  └─ change_requests/            # drop CR memos here for updates
└─ output/                        # ▶ OneData creates these
   ├─ requirements/               # BRD, KPIs, KBQs
   ├─ design/                     # ARD/RRD, pipeline design, models
   ├─ sttm/                       # mappings (CSV + Excel)
   ├─ generated/                  # DDL, DML, scores
   ├─ versions/                   # immutable snapshots
   └─ logs/
📤

You only touch input/. Drop your requirements and metadata in, and the agents create and update everything under output/ — phase by phase, with approval gates and validation between each.

Four engagement scenarios

OneData detects what you've provided and starts at the right phase — you don't always begin from scratch

ScenarioYou providePipeline starts at
1Raw evidence only (transcripts, docs, SQL)Requirements Generation Agent → Design → …
2Raw evidence + draft ARD/RRD designsDesign validates your drafts, then continues
3Raw evidence + designs + pipeline planDesign fills the table-level detail, then continues
4Full inputs (BRD + models + ER)Straight to STTM Generation Agent → Code → dbt
🔍

Scenario detection is automatic — a quick detect-scenario check reads your input/ folder and picks the right entry point.

What happens, phase by phase

Each phase does its work, asks for what it needs, and produces a reviewable artifact before the next begins

  1. Launch & set upOpen the project in VS Code and start Cortex Code from the root. It discovers the .cortex package, creates the output folders, and reads your inputs.
  2. Requirements & designThe agents draft business requirements and the pipeline blueprint (data layers, tables, relationships), and ask you to approve before moving on.
  3. Choose the mapping styleYou pick Direct or Layered — layered adds reusable staging/intermediate tables when the logic is complex and shared across outputs.
  4. Approve the designBefore any mappings are written, you review the full table plan and dependency graph, and approve or request changes.
  5. Generate the mappings (STTM)The agent writes the source-to-target mapping for every table and hands you a review-ready Excel workbook.
  6. Independent validation & sign-offA read-only validator checks the mappings against the requirements and models and returns a clear verdict. The reviewed sheet becomes the formal contract — code generation cannot start without your sign-off.
  7. Generate the codeYou choose SQL or Snowpark. The agent builds the code in dependency order, and an independent check scores each piece against the specification.
  8. Package into dbtEverything is assembled into a ready-to-run dbt project — models, reusable logic, business marts, and quality tests — then verified end to end.
  9. Version & hand offThe completed run is snapshotted and archived. From here, later changes go through the Incremental Updates/Enhancements Agent agent instead of a full rebuild.

The commands, in short

Everything runs through one entry point, pipeline_cli.py. A typical run looks like this — the agents drive it, so you mostly approve gates.

Terminal — typical run
# Detect what you provided and set up the workspace
python scripts/pipeline_cli.py extract-inputs
python scripts/pipeline_cli.py detect-scenario

# Requirements & design run as agents, with approval gates …
# then, from mappings onward:
python scripts/pipeline_cli.py init
python scripts/pipeline_cli.py discover
python scripts/pipeline_cli.py validate-sttm      # independent check
python scripts/pipeline_cli.py excel              # review-ready workbook

# After sign-off: generate, score, package
python scripts/pipeline_cli.py generate-ddl <ENTITY>
python scripts/pipeline_cli.py score <ENTITY>
python scripts/pipeline_cli.py dbt

# Version the completed run for safe incremental updates
python scripts/pipeline_cli.py snapshot
python scripts/pipeline_cli.py archive