How extraction works
When you create a table, Lasso processes your source files through an AI pipeline:- Document parsing — PDFs, spreadsheets, and images are parsed to extract raw text and visual elements.
- Schema mapping — The AI maps the raw content to your schema’s column definitions.
- Row generation — Each detected product or item becomes a row with values for each column.
- Validation — Extracted values are validated against column types (numbers, URLs, emails, etc.).
Source types
You can provide data to extract in two supported ways:
For a remote file, download it and upload it through
POST /files before creating the table. The extraction worker does not support direct file_urls ingestion.
Table lifecycle
A table moves through these statuses:- queued or queued_for_processing — The job is waiting for a worker.
- processing or processing_by_worker — Extraction is actively running. The
progressfield (0-100) tracks completion. - completed — All rows have been extracted and are ready to query.
- error — Something went wrong. Check
error_messagefor details. - cancelled — Processing was cancelled.
Polling vs webhooks
You have two options to know when extraction finishes:- Polling — Poll
GET /tables/{table_id}until status iscompleted,error, orcancelled. - Webhooks — Pass a
webhook_urlwhen creating the table. Lasso sends an HTTP POST when processing completes.
Working with rows
Once a table is completed, each extracted item is a row. Rows contain:- data — A key-value object matching your schema columns.
- validation_status — Whether the row passed type validation.
- enhancement_status — Per-column status of any AI enhancements.
- is_edited — Whether the row was manually modified via the API.

