clgraph

Context layer for warehouse AI

clgraph — the context layer between your warehouse SQL and any LLM.

clgraph parses every statement into one column-level graph: tables, CTEs, subqueries, and how each column derives from the last. Serve that graph to a model through Python, a CLI, or an MCP server, and it stops guessing about your data — static analysis only, no database connection.

What the graph gives a model

LLM column descriptions

Descriptions are generated lineage-aware: a column's upstream sources and transformations feed the prompt, so docs stay grounded in the SQL.

pipeline.generate_all_descriptions()

Text-to-SQL

Ask a question in plain English. clgraph hands the model your pipeline's real schema, labels each table source, intermediate or final, and adds join hints taken from the joins your own SQL already performs — then Studio shows you that context beside the SQL it produced.

"Which customers have generated the most lifetime revenue?"

How the graph gets built

Step 1

Parse

Every statement is parsed into a graph. Each column — in tables, CTEs, subqueries — becomes a node; each expression becomes an edge.

Step 2

Combine

Graphs from separate files concatenate: where one query writes a table another reads, the graphs join.

dbt ref() and {{templated}} names are just flavors of SQL input — they resolve into the same graph.

Step 3

Complete graph

Feed it every SQL file in your warehouse and you get the complete column-level lineage graph.

Step 4

Graph as LLM context

The model's context comes out of this graph: tables, columns, and the derivation paths behind them — not a hand-maintained data dictionary.

Pure static analysis — clgraph never connects to your database or reads execution logs.

Open Studio →See everything it does →