Skip to main content
Tables are the core resource in Lasso. A table represents an extraction job that takes unstructured source data and produces structured rows of product information.

How extraction works

When you create a table, Lasso processes your source files through an AI pipeline:
  1. Document parsing — PDFs, spreadsheets, and images are parsed to extract raw text and visual elements.
  2. Schema mapping — The AI maps the raw content to your schema’s column definitions.
  3. Row generation — Each detected product or item becomes a row with values for each column.
  4. Validation — Extracted values are validated against column types (numbers, URLs, emails, etc.).

Source types

You can provide data to extract in two supported ways: For a remote file, download it and upload it through POST /files before creating the table. The extraction worker does not support direct file_urls ingestion.

Table lifecycle

A table moves through these statuses:
  • queued or queued_for_processing — The job is waiting for a worker.
  • processing or processing_by_worker — Extraction is actively running. The progress field (0-100) tracks completion.
  • completed — All rows have been extracted and are ready to query.
  • error — Something went wrong. Check error_message for details.
  • cancelled — Processing was cancelled.

Polling vs webhooks

You have two options to know when extraction finishes:
  • Polling — Poll GET /tables/{table_id} until status is completed, error, or cancelled.
  • Webhooks — Pass a webhook_url when creating the table. Lasso sends an HTTP POST when processing completes.

Working with rows

Once a table is completed, each extracted item is a row. Rows contain:
  • data — A key-value object matching your schema columns.
  • validation_status — Whether the row passed type validation.
  • enhancement_status — Per-column status of any AI enhancements.
  • is_edited — Whether the row was manually modified via the API.
Rows can be updated individually, in bulk, or deleted. See Rows API for details.