
Capture documents and operational data.
Read scans, PDFs, spreadsheets, emails and system exports into one place, without defining a field list first. Every value keeps the line it came from.
AI agents can’t run a business until someone structures both the information and the business rules behind it. Talonic turns documents and operational data into a reusable Data Registry, so you can access the data you need and automate the work your company wants to.
Ask the document estate
↳ ▋
| counterparty | value | renews | source · confidence |
|---|---|---|---|
| Northern Utilities | €1.24M | Sep 04 | p.12 · 0.96 |
| Municipal Heat Ltd | €640K | Sep 11 | p.08 · 0.94 |
| East Industrial Park | €2.08M | Sep 13 | p.21 · 0.97 |
3 rows in 0.38 s. Every value traces to its source line.
Trusted across logistics, energy, automotive and contract operations.

Read scans, PDFs, spreadsheets, emails and system exports into one place, without defining a field list first. Every value keeps the line it came from.

The same fields, entities and relationships resolve across every document you add. People and AI agents query the Registry, compare records and trace answers back to their sources.

Mine operational changes for candidate business rules. Validate them with your experts, apply approved rules to repeatable work, and route exceptions to a person.

Reads every contract before a person does and holds the ones a rule flags.
How Contract Review works
Checks that freight documents agree and clears loads for billing, beside the biller.
How Auto Billing works
Finds the money owed where agreed terms and actual transactions differ.
How Money Found worksChoose which Registry data an agent can access. It can compare records and investigate exceptions, with source evidence and logged activity.
Cases that fail the approved conditions go to a person. Review those decisions to propose and validate additional rules before activation.

Read the contract family once. Reuse it across departments.
The migration needed 75 fields per contract dataset. One read of the estate holds 1,042,665 source observations and 52,937 reusable concepts, live tenant, 26 August 2026.
Live tenant since 26 August 2026.
Read the case study
Match freight documents to loads. Test billing decisions alongside the team.
The same hand labelled documents, two fields, both systems scored the same way. Replacing a six figure incumbent.
In production. Auto Billing in shadow run since 12 August 2026.
Read the case study
Read printed curves once. Reuse them as engineering data.
Performance and reliability data that existed only as printed curves on technical datasheets. Talonic finds each chart, reads its axes and curves, and delivers the X/Y series as tables an engineer can reuse: 736 charts, about 3,300 curves, each sampled to 100 points and checked back against the page.
Method and pilot established. Detected charts delivered, scale-up follows approval.
Read the case study
A multilingual contract estate across thirty countries, in formats that never agreed, that purchasing could not query. One batch read once, text and image alike, and delivered as one row per contract in the shape Ivalua imports. The programme runs to 22,000 contracts.
Pilot completed. Batch delivered and received, ERP import-ready.
Read the case study
Management reports that were already approved, already read, and already filed. Nineteen months from 24 verticals, with drifting labels and formats normalised into one queryable layer. Nobody had to rebuild the systems underneath.
PoC and insight package completed. Scale-up proposed.
Read the case study
German school timetable law from 1996 to 2021, transcribed exactly as printed. Every row carries its source page, a confidence score, and a note on anything ambiguous.
One batch delivered. The corpus gated.
Read the case studyRead documents and operational records into the Registry without defining a field list first. Preserve the source behind each value.
Let people and AI agents ask questions across authorized Registry data, not just one document or one report.
Compare database snapshots. Agents investigate changes alongside source documents and propose rules behind recurring decisions.
Your experts approve the rules. Apply them to eligible cases and route unresolved exceptions for review.

Most tools return a value. Talonic returns a value, the line it came from, the region of the scan that produced that line, the confidence, the phase, and the reasoning. Auditable by default. Defensible by design.
One result per benchmark. The protocol, the corpus and the rows we lose are one click away.
Structure-first answered 5 of 5 headline ranking queries on 28 SEC 10-K filings with XBRL gold labels, under a preregistered protocol. BM25 top-12 retrieval answered 0 of 5. The 39-query battery over 53 filings, including the rows Talonic loses, is on the benchmark page.
Read the benchmarkReading the 53-filing corpus once cost $10.14. Structured queries on it then add no retrieval cost per question, because nothing is re-read. The agent path does add cost per question, and its line is on the same page.
Read the benchmarkThe same 100 questions, asked ten times: the structured path returned byte-identical answers on every pass. Hybrid RAG on the same corpus matched its own earlier answers on 2 of 39 questions.
Read the benchmarkREST API, Node SDK, or MCP server (Model Context Protocol). Every response carries typed fields, confidence, and provenance, so your agent knows when not to trust an answer. Free tier: 5,000 credits a month, no credit card.
npx -y @talonic/mcp@latestWithout a registry, adding a downstream system means re-running every extraction, so cost scales as documents times systems, and every new agent, workflow or integration multiplies it. Retrieval has the same shape: it re-reads the corpus on every question, re-spends the tokens, and can return a different answer tomorrow.
Talonic reads each document once, resolves the schema from the documents themselves, and keeps a registry where every value traces to its source line, region, confidence and reasoning. Every new system queries that registry instead of extracting again.
See the published benchmark against RAGDOCUMENT ONTOLOGY
From Schedule K-1 to Bill of Lading (Ocean), from Notarial Deeds to QC Inspection Forms. Talonic understands the structure of every enterprise document type out of the box.
Schedule K-1, Form 1099-MISC, Form W-8BEN, Form 1040, Form 990, and 48 more
Purchase Order, Commercial Invoice, Pro Forma Invoice, Goods Receipt Note, Three-Way Match Report, and 48 more
Bill of Lading (Ocean), Bill of Lading (Inland), Air Waybill (AWB), House Air Waybill (HAWB), Sea Waybill, and 48 more
Notarial Deed, Master Service Agreement, Statement of Work, Non-Disclosure Agreement, License Agreement, and 48 more
Commercial Register Extract, Certificate of Good Standing, Certificate of Incorporation, Annual Return, Director's Report, and 48 more
Patient Intake Form, Medical History Questionnaire, Informed Consent Form, Discharge Summary, Operative Report, and 48 more
QC Inspection Form, Certificate of Analysis (CoA), Certificate of Conformance, First Article Inspection Report, Material Test Report, and 48 more
Insurance Application, Policy Declaration Page, Insurance Policy Wording, Endorsement / Rider, Certificate of Insurance, and 48 more
Property Deed, Title Report, Land Registry Extract, Survey Plan, Zoning Certificate, and 48 more
Employment Application, Offer Letter, Employment Contract, Background Check Report, I-9 Employment Eligibility, and 47 more
DIN SPEC 91491

Talonic co-authored DIN SPEC 91491 with Fraunhofer IIS, Humboldt-Innovation, GIIC, and the German standards body. Europe's first standard for AI-ready data at the schema layer.
DIN SPEC 91491 defines what schema-layer data readiness looks like. Enterprises aligning with the standard need an implementation that was built alongside it.
Talonic is a reusable data registry for difficult enterprise information. It captures documents once, discovers and organizes the fields, entities and relationships it finds with provenance, and makes that data reusable across workflows, systems and agents. No schema or field list is required before capture; downstream consumers map the captured data into the output they need.
Talonic supports over 25 file formats including PDFs, scans, Word documents, Excel spreadsheets, images, emails, and more, classified into a 529-type document ontology.
Yes. Talonic is GDPR and HIPAA compliant, ISO 27001 and ISO 42001 aligned, and co-authored DIN SPEC 91491. All data is processed in EU-resident infrastructure in Germany West Central.
Europe's first standard for AI-ready data at the schema layer, published in 2025. Talonic co-authored the standard alongside Fraunhofer IIS, Humboldt-Innovation, and GIIC.
OCR converts pixels to text. Talonic classifies documents against a 529-type ontology, extracts schema-validated fields with confidence scores, reconciles entities into cases, and delivers typed data to systems of record with full audit trail.
Yes. Send a representative sample of your documents and Talonic runs them within five business days. You receive a field-level analysis of extractability, confidence distribution, recommended schema, and a production roadmap.
Send a representative sample of contracts, scans, case files, or operational documents. We run them through Talonic and send back your extracted data within five business days: field-level coverage, confidence, and provenance. If the database doesn’t beat what you have, you’ll know exactly why, field by field.