clgraph

What clgraph does

LLM column descriptions

Descriptions are generated lineage-aware: a column's upstream sources and transformations feed the prompt, so docs stay grounded in the SQL.

pipeline.generate_all_descriptions()

Text-to-SQL

Ask a question in plain English. clgraph hands the model your pipeline's real schema, labels each table source, intermediate or final, and adds join hints taken from the joins your own SQL already performs — then Studio shows you that context beside the SQL it produced.

"Which customers have generated the most lifetime revenue?"

Model Context Protocol server

Point any MCP client at your pipeline and the graph becomes callable: tools for lineage (trace_backward, get_lineage_path), schema (search_columns, get_relationships) and governance (find_pii_columns), plus the full schema and table list as MCP resources. No glue code.

python -m clgraph.mcp --pipeline path/to/queries/

Column-level lineage

Every column in every table, CTE, and mart becomes a node. Click a column in the graph and Studio runs the same trace the library exposes in Python.

pipeline.trace_column_backward(
    "mart_customer_ltv", "lifetime_revenue"
)

Any flavor of SQL input

Warehouse SQL rarely comes plain: table names are {{templated}} from config, dbt models reference each other with ref(). To clgraph these are just flavors of the same input — they resolve into one graph.

SELECT o.order_id, c.customer_name
FROM {{ ref('stg_orders') }} o
JOIN {{ ref('stg_customers') }} c USING (customer_id)

clgraph also does