Agents write specifications. A deterministic engine executes them.
AI can now write your pipelines, but "English in, pipeline out" magic is how trust dies in production. Starflow takes a different bet: everything an agent produces is schema-validated, git-reviewable YAML and SQL. Everything that runs is deterministic engine code.
Starflow is the Declarative Data StackΒ where humans and AI agents author the same specs, under the same guardrails, and the LLM never enters the execution path.
Declarative Data Stack
Unified, simple, and powerful data engineering
See it in action
Three ways to ship a pipeline
Talk to a Starflow persona, ask Starflow what to do next, or call a CLI skill directly. Every interaction below is the real prompt and the real shape of the response.
Talk to a Starflow persona
Five expert AI personas cover the data lifecycle. Winston the Data Architect lays out two or three options with trade-offs before naming a recommendation, and names the failure mode of every choice.
Meet the personas βAsk Starflow what to do next
Starflow reads a manifest of skill dependencies and scans the artifacts you've already produced. It recommends the next required step based on real state, not a fixed checklist, so it works whether you started at discovery or jumped in mid-stream.
See the manifest model βOr call a CLI skill directly
For quick, targeted tasks, skip the methodology. Each of the 53 CLI skills knows the exact flags, YAML schema, and write strategies for its Starflow command. Ask in plain English; receive correct configuration on the first try.
Browse the catalog βDeveloper Tools
Agent-Native Data Engineering
Starflow meets you, and your agents, where you work. An MCP server that exposes the engine as tools, 74 AI skills for Claude Code, Copilot, and Gemini, and a VSCode extension for humans. Context out, feedback in, guardrails around.
MCP Server
The engine as a tool surface for agents
Point any MCP client at the Starflow engine and it gets validated, lineage-aware, cost-estimated tool responses, not log prose. Every error is a structured object with a code, file, YAML path, and remediation hint, so an agent self-corrects in one turn instead of burning ten.
Validate & Infer
Schema-checked YAML, structured JSON errors with fix hints
Lineage Tools
lineage, col-lineage, table-dependencies as tool calls
Dry-Run with Cost
Agents see the price of a transform before it runs
Tests & Expectations
Machine-readable results close the verification loop
Freshness & Audit
Structured events an agent runbook can act on
Guardrail Policies
Required tests, cost ceilings, protected domains, enforced on every author
VSCode Extension
Your entire data platform, inside your editor
Build, test, and deploy data pipelines without leaving VSCode. Auto-infer schemas, preview SQL, visualize lineage, and deploy DAGs with a single click.
Schema Inference
Auto-generate YAML configs from data sources
SQL Preview & Dry Run
Validate transformations before execution
Visual Lineage & ER Diagrams
Understand data flow at a glance
One-Click DAG Deploy
Generate and deploy Airflow/Dagster workflows
Results Viewer
Inspect query results inside the editor
ACL Management
Review and manage access control visually
AI Skills
74 skills that make any coding agent a Starflow expert
Install the Starflow Skills plugin and your agent instantly knows every CLI command, YAML pattern, write strategy, and best practice: 53 CLI skills plus 21 personas and workflows, verified in CI against a real engine build on every release.
Ingestion Skills
autoload, load, ingest, stage, kafkaload and more
Transform & Extract
transform, extract, extract-data, extract-schema
Lineage & Quality
lineage, col-lineage, expectations, table-deps
Ops & Orchestration
dag-generate, dag-deploy, serve, metrics, freshness
Schema & Security
bootstrap, infer-schema, secure, iam-policies
Personas & Workflows
Winston the Data Architect and 20 more expert workflows
The Old Way vs. The Starflow Way
See the dramatic difference in approach, complexity, and results
Compare the fragmented Modern Data Stack approach with Starflow's unified Declarative Data Stack solution.
The Old Way
Modern Data Stack complexity
Fragmented Quality
Separate tools for ingestion, quality, and transformation create gaps where bad data slips through.
Template Lock in
SQL tangled in Jinja/Python templates that only work in specific tools.
Expensive Development
Every change requires expensive warehouse compute cycles for testing.
Tool Sprawl
Multiple separate tools for each function, creating integration complexity.
Agents Without Guardrails
Chat-generated SQL and Python pasted into production with no validation, no lineage, no review trail.
The Problem
The Modern Data Stack approach creates complexity, vendor lock in, and expensive development cycles that slow down your team and increase costs.
Ready to Build the Future of Data? π
Give your data team - human and agent alike - one set of specs, one set of guardrails, one deterministic engine.
Join the teams already building with Starflow's Declarative Data Stack approach.
Start Your Journey Today
Experience the power of the Declarative Data Stack. See how Starflow can transform your data engineering workflow in just 30 minutes.
Free to get started. Open source, no vendor lock in, and works with your existing infrastructure.