Starflow

Agents write specifications. A deterministic engine executes them.

AI can now write your pipelines, but "English in, pipeline out" magic is how trust dies in production. Starflow takes a different bet: everything an agent produces is schema-validated, git-reviewable YAML and SQL. Everything that runs is deterministic engine code.

Starflow is the Declarative Data StackΒ where humans and AI agents author the same specs, under the same guardrails, and the LLM never enters the execution path.

Starflow
=
πŸš€

Declarative Data Stack

Unified, simple, and powerful data engineering

βœ… Quality-First
Built in validation
⛓️ Pure SQL
No templating
πŸ’» Local Dev
Free & fast

See it in action

Three ways to ship a pipeline

Talk to a Starflow persona, ask Starflow what to do next, or call a CLI skill directly. Every interaction below is the real prompt and the real shape of the response.

Talk to a Starflow persona

Five expert AI personas cover the data lifecycle. Winston the Data Architect lays out two or three options with trade-offs before naming a recommendation, and names the failure mode of every choice.

Meet the personas β†’
~/projects/acme-analytics - claude code
β€Ί

Ask Starflow what to do next

Starflow reads a manifest of skill dependencies and scans the artifacts you've already produced. It recommends the next required step based on real state, not a fixed checklist, so it works whether you started at discovery or jumped in mid-stream.

See the manifest model β†’
~/projects/acme-analytics - claude code
β€Ί

Or call a CLI skill directly

For quick, targeted tasks, skip the methodology. Each of the 53 CLI skills knows the exact flags, YAML schema, and write strategies for its Starflow command. Ask in plain English; receive correct configuration on the first try.

Browse the catalog β†’
~/projects/acme-analytics - claude code
β€Ί

Developer Tools

Agent-Native Data Engineering

Starflow meets you, and your agents, where you work. An MCP server that exposes the engine as tools, 74 AI skills for Claude Code, Copilot, and Gemini, and a VSCode extension for humans. Context out, feedback in, guardrails around.

MCP Server

The engine as a tool surface for agents

Point any MCP client at the Starflow engine and it gets validated, lineage-aware, cost-estimated tool responses, not log prose. Every error is a structured object with a code, file, YAML path, and remediation hint, so an agent self-corrects in one turn instead of burning ten.

Validate & Infer

Schema-checked YAML, structured JSON errors with fix hints

Lineage Tools

lineage, col-lineage, table-dependencies as tool calls

Dry-Run with Cost

Agents see the price of a transform before it runs

Tests & Expectations

Machine-readable results close the verification loop

Freshness & Audit

Structured events an agent runbook can act on

Guardrail Policies

Required tests, cost ceilings, protected domains, enforced on every author

MCPClaude CodeCopilotGeminiDuckDBBigQuerySnowflakeRedshift

VSCode Extension

Your entire data platform, inside your editor

Build, test, and deploy data pipelines without leaving VSCode. Auto-infer schemas, preview SQL, visualize lineage, and deploy DAGs with a single click.

Schema Inference

Auto-generate YAML configs from data sources

SQL Preview & Dry Run

Validate transformations before execution

Visual Lineage & ER Diagrams

Understand data flow at a glance

One-Click DAG Deploy

Generate and deploy Airflow/Dagster workflows

Results Viewer

Inspect query results inside the editor

ACL Management

Review and manage access control visually

BigQuerySnowflakeRedshiftDatabricksDuckDB

AI Skills

74 skills that make any coding agent a Starflow expert

Install the Starflow Skills plugin and your agent instantly knows every CLI command, YAML pattern, write strategy, and best practice: 53 CLI skills plus 21 personas and workflows, verified in CI against a real engine build on every release.

Ingestion Skills

autoload, load, ingest, stage, kafkaload and more

Transform & Extract

transform, extract, extract-data, extract-schema

Lineage & Quality

lineage, col-lineage, expectations, table-deps

Ops & Orchestration

dag-generate, dag-deploy, serve, metrics, freshness

Schema & Security

bootstrap, infer-schema, secure, iam-policies

Personas & Workflows

Winston the Data Architect and 20 more expert workflows

Claude CodeCopilotGeminiDuckDBBigQuerySnowflakeRedshiftDatabricks

The Old Way vs. The Starflow Way

See the dramatic difference in approach, complexity, and results

Compare the fragmented Modern Data Stack approach with Starflow's unified Declarative Data Stack solution.

The Old Way

Modern Data Stack complexity

Fragmented Quality

Separate tools for ingestion, quality, and transformation create gaps where bad data slips through.

Quality checks as afterthought
Bad data reaches production
Downstream debugging nightmares
Multiple tool dependencies

Template Lock in

SQL tangled in Jinja/Python templates that only work in specific tools.

Vendor specific templating
Impossible to test locally
No portability between tools
Complex debugging required

Expensive Development

Every change requires expensive warehouse compute cycles for testing.

Expensive warehouse compute
Slow feedback loops
Wasted development time
High operational costs

Tool Sprawl

Multiple separate tools for each function, creating integration complexity.

Multiple vendor dependencies
Complex integration points
Fragile system architecture
Maintenance overhead

Agents Without Guardrails

Chat-generated SQL and Python pasted into production with no validation, no lineage, no review trail.

LLM output runs directly
No schema validation
Unreviewable generated code
Trust erodes with every incident

The Problem

The Modern Data Stack approach creates complexity, vendor lock in, and expensive development cycles that slow down your team and increase costs.

Ready to Build the Future of Data? πŸš€

Give your data team - human and agent alike - one set of specs, one set of guardrails, one deterministic engine.

Join the teams already building with Starflow's Declarative Data Stack approach.

Faster development cycles
Lower operational costs
Simplified architecture
Better data quality
Team productivity
Vendor independence
Agent-ready by design
No LLM at runtime

Start Your Journey Today

Experience the power of the Declarative Data Stack. See how Starflow can transform your data engineering workflow in just 30 minutes.

Free to get started. Open source, no vendor lock in, and works with your existing infrastructure.

Trusted by Data Teams Worldwide

🏒 Enterprise Ready
πŸ”’ Production Grade
⚑ High Performance
🌍 Open Source