Guides

Data work, run like production.

Seven guides, one lens: a model does the semantic work, a server keeps the state, and every accepted value has an acceptance record behind it. They stand on their own — read them with or without the product.

Start here

AI data pipeline: what it is and how to build one that survives the session

The five stages, the state question that separates pipelines from demos, and a seven-box checklist you can argue with.

Read the guide →
Choosing tools

Data extraction tools in 2026: parsers, models, and the accounting in between

Four tool families sorted by input, a decision path, and the acceptance accounting every comparison skips.

Read the guide →
Core mechanism

LLM data extraction: batches, gates, and why attempt #2 exists

Four failure modes, pull-based batches sized to judgement, the golden gate's three modes, and the ledger where retries are instruments.

Read the guide →
The contract

Structured data extraction with LLMs: schema first, then scale

Valid JSON is not valid data. What constraint tooling guarantees, what it can't, and the declared checks that judge every submission.

Read the guide →
Growing columns

Data enrichment: growing the columns your data never had

Lookup versus judgement, the 2026 tool landscape sorted honestly, and why enriched is not the same as trusted.

Read the guide →
The clean edge

AI data cleaning: mechanical rules first, model judgement second

Where rules end and models begin, five patterns for a model in the loop, and the account every rejected row deserves.

Read the guide →
Collection

AI web scraping: from brittle selectors to agent-run collection

Three modes wearing one name, where dumb selectors still win, and why most scraped datasets die in stage two.

Read the guide →
The full argument

How Tablize works — the mechanism, argued in full

Two dead ends, four steps, three gate modes, and the honest boundary of what a landed row does and does not mean.

Read the argument →

Public beta · free · macOS, Linux and Windows

Grow the next column your data never had.