Skip to main content

THE REGISTRY LAYER FOR UNSTRUCTURED DATA

Ingest once.
Query forever.

Talonic ingests messy enterprise documents once — PDFs, scans, handwriting, spreadsheets, emails — into persistent, schema-validated records with field-level evidence. Every workflow, system, and AI agent queries the same registry later. The source is never re-extracted.

Response within 1 business day · Or start free: 5,000 credits/mo, no credit card

Co-author of DIN SPEC 91491, Europe’s first AI data standard · GDPR · HIPAA · ISO 27001 / 42001 aligned · EU-resident infrastructure

Talonic dashboard — 90-document searchable library, 11,550 extracted knowledge observations resolving into 9,393 reusable concepts, with knowledge-compression and collection-shape views

TRUSTED ACROSS LOGISTICS, ENERGY, AUTOMOTIVE, AND CONTRACT OPERATIONS

Maruti Suzuki
GETEC
Phoenix
Bridgeway
Ricano
DIN
WZB
DSR
Performing Digital
Consulting SIM
Maruti Suzuki
GETEC
Phoenix
Bridgeway
Ricano
DIN
WZB
DSR
Performing Digital
Consulting SIM

THE PROBLEM

The data was never made reusable.

Most tools extract for one schema, then stop. The structure is thrown away with the workflow that asked for it — and the next system, report, or agent starts again from the raw documents.

Every new system re-extracts everything.

Without a registry, adding a downstream system means re-running every extraction. Cost scales as documents times systems.

Every new agent, workflow, or integration multiplies your extraction cost. Unless the data is already in a registry.

12345678910Downstream systems / agentsTotal extraction workWithout a registryO(n × m)With TalonicO(n) + O(m)5×
5 5×

The same root cause, measured four ways:

CAPTUREDIDC · Data Age 2025

10%

Of enterprise unstructured data is ever analyzed.

STRUCTUREDGartner · 2024

5%

Of enterprise data is queryable by AI today.

CONNECTEDMuleSoft · 2024 Connectivity Benchmark

33%

Of enterprise applications are integrated.

REUSABLEMIT NANDA · The GenAI Divide, 2025

5%

Of enterprise GenAI pilots reach production.

Four broken layers. None of them fixed by a better model. As long as extraction is disposable, every use case pays the O(n × m) bill.

REUSE, NOT RE-EXTRACTION

Solve the workflow. Keep the data.

Run the document work you already pay for. Every field Talonic touches lands in the registry, typed and queryable, so the next agent, schema, or report is served from structure instead of a fresh extraction. The platform is the byproduct of getting the work done.

  • ·Vendor and customer contracts
  • ·AP/AR: invoices, POs, receipts
  • ·Logistics: manifests, bills of lading
  • ·Claims, KYC, and onboarding files

LIVE DEPLOYMENTS

Live in regulated enterprise. Today.

GETECEnergy · Germany

8,500 active contracts under structuring as of Q2 2026. Schema v2, 59 German-language fields, Microsoft Dynamics injection in Q2 2026.

Next: automated anomaly detection across the full contract portfolio.

Phoenix GroupPharma · Germany

22,000 vendor contracts structuring to Ivalua. Commercial execution Q2 2026.

Next: extending schema coverage to clinical trial documentation.

BridgewayLogistics · USA

930-document ground-truth benchmark. Accuracy from 75% to 92% across POC cycles (2025–2026). Replacing a six-figure incumbent system.

Next: full carrier-to-load matching across the Gemspring portfolio.

30,500+

enterprise contracts under structuring today.

6-figure

incumbent system being replaced at Bridgeway.

< 200ms

median per-document processing time.

THE MECHANISM

How Talonic works.

How Talonic works

01 Sources

An endless stream of documents. None of it queryable. Ingested through Talonic from email, drive, S3, SFTP, or API. Document types include contracts, invoices, emails, scans, carrier manifests, purchase orders, service reports, and allocations.

02 Data Capture

Every document, parsed. Every field, extracted. Every schema, recognized. In parallel, at scale. Talonic processes service agreements, invoices, carrier manifests, purchase orders, and service reports — extracting 64 to 113 fields per document with schema recognition (vendor_contract_v7, invoice_v3, carrier_manifest_v2, purchase_order_v2, service_report_v1).

03 Field Registry

Every field becomes canonical. Every extraction compounds. The registry never resets. Canonical fields include vendor_name, contract_value, effective_date, term_end, auto_renew, governing_law, notice_period, invoice_number, total_amount, tax_rate, currency, due_date, payment_terms, load_id, carrier, origin, destination, weight_kg, incoterms, vehicle_type, technician, report_id, service_date, and facility. New canonicals are continuously discovered: meter_id, incoterm_variant, clinical_phase, hs_code, vat_id, hazmat_class, and more.

04 Schema Matching

Every new schema, mapped to canonicals. Every field, recognized regardless of source terminology. Continuously. Example: customer_contract_v3 maps party_name to vendor_name (0.94 confidence), agreement_value to contract_value (0.91), termination_date to term_end (0.89), monetary_unit to currency (0.97). 4/4 matched with 0.93 average confidence.

05 Query

Every question, answered against the registry. Every answer, structured and typed. Continuously. Query with SQL, natural language, or API calls. Examples: "SELECT contract_value FROM canonicals WHERE governing_law = BGB § 305ff", "which contracts expire in Q4 2026?", "talonic.query(canonicals, filter={governing_law: BGB § 305ff})". Deliver to SAP S/4HANA, Salesforce CRM, NetSuite, Dynamics, Ivalua, or any REST endpoint.

Most document AI extracts for a schema. Talonic builds the data layer behind every future schema.

Traditional parsing tools start with a destination format and extract into it. Extraction quality is the floor, not the ceiling. Customers consistently see 90%+ accuracy in head-to-head benchmarks against incumbents. And with Talonic, that result lives in a registry that compounds across every future schema, system, and agent.

PROVENANCE

Every value points back to where it came from.

Most tools return a value. Talonic returns a value, the line it came from, the region of the scan that produced that line, the confidence, the phase, and the reasoning. Auditable by default. Defensible by design.

EINGEGANGEN14.01.2026
→ check term
EINGEGANGEN
14.01.2026
 
 
SERVICE AGREEMENT
 
This Service Agreement ("Agreement") is entered
into on 01. January 2026 between:
 
Meridian Energy AG, a corporation organized
under the laws of the Federal Republic of
Germany, with its principal place of
business at Friedrichstraße 200, 10117
Berlin ("Supplier"); and
 
Kent Logistics Ltd., a corporation organized
under the laws of England and Wales, with
its principal place of business at 14
Cannon Street, London EC4N 6JJ ("Customer").
 
1. TERM
 
This Agreement shall commence on 01.01.2026
and shall remain in effect until 31.12.2027
unless terminated earlier in accordance
with Section 8.
 
2. CONSIDERATION
 
The total contract value payable by Customer
to Supplier under this Agreement shall be
EUR 2.480.000 (two million four hundred
eighty thousand Euros), payable in
accordance with the payment schedule in
Exhibit A.
 
3. AUTO-RENEWAL
 
Upon expiration of the initial term, this
Agreement shall automatically renew for
successive one-year terms unless either
party provides written notice of non-renewal
at least ninety (90) days prior to the
then-current expiration date.
[page 1 of 2]
~ M. Richter
4. GOVERNING LAW
 
This Agreement shall be governed by and
construed in accordance with the laws of the
Federal Republic of Germany, specifically
BGB § 305ff.
 
5. CURRENCY AND PAYMENT
 
All payments shall be made in Euros (EUR)
to the account designated by Supplier in
writing.
 
6. DELIVERY POINT
 
Services shall be delivered at Berlin HKW,
Heizkraftwerk Mitte.
 
7. VOLUME
 
Supplier shall deliver no less than 14,400
MWh per annum.
 
IN WITNESS WHEREOF, the parties have
executed this Agreement as of the date
first written above.
 
 
________________________________
For Meridian Energy AG
 
 
________________________________
For Kent Logistics Ltd.
[page 2 of 2]
Service Agreement
Parties
Supplier: Meridian Energy AG, incorporated under the laws of
the Federal Republic of Germany. Principal place of business:
Friedrichstraße 200, 10117 Berlin.
Customer: Kent Logistics Ltd., incorporated under the laws of
England and Wales. Principal place of business: 14 Cannon Street,
London EC4N 6JJ.
Term
This Agreement runs from 01.01.2026 to 31.12.2027, unless
terminated earlier per Section 8.
Consideration
Total contract value: EUR 2,480,000 (two million four hundred
eighty thousand Euros).
Auto-Renewal
Automatic renewal for successive one-year terms unless written
notice of non-renewal is given at least 90 days prior to the
then-current expiration date.
Governing Law
Federal Republic of Germany. BGB § 305ff applies.
Currency
EUR. All payments in Euros to Supplier's designated account.
Delivery Point
Berlin HKW (Heizkraftwerk Mitte).
Volume
Not less than 14,400 MWh per annum.

Execution: signed by both parties on the date of the Agreement.
FIELDVALUECONF
  • vendor_nameMeridian Energy AG0.99
  • customer_nameKent Logistics Ltd.0.97
  • contract_start_date2026-01-010.99
  • contract_end_date2027-12-310.99
  • contract_value2,480,0000.97
  • currencyEUR1.00
  • auto_renewtrue0.96
  • notice_period_days900.94
  • governing_lawDE · BGB § 305ff0.92
  • delivery_pointBerlin HKW0.93
  • annual_volume_mwh144000.95
  • schema_versionvendor_contract_v21.00
Click any field to trace it back to the document.

USE CASES

One registry. Every document workflow.

The same ingested corpus serves every downstream consumer. These are the four we see most.

Accounts payable

Invoices, POs, receipts, and credit notes land as typed records. Three-way match runs against the registry — not against PDFs in a shared drive.

Commercial Invoice · Purchase Order · Goods Receipt Note

Invoice parsing →
Contract operations

Parties, terms, renewal windows, and obligations extracted once, tracked across amendments and side letters. Cases group what belongs together.

Master Service Agreement · Contract Amendment · Side Letter

How cases work →
AI agents

Point an agent at the registry via MCP or REST and it reads any document like a typed database: fields, confidence, and provenance on every response.

MCP server · REST API · Node SDK

For agents →
Context layer

Serve copilots and internal tools the same governed fields your ERP consumes. One extraction, every consumer — no parallel pipelines to keep in sync.

Dynamics 365 · Ivalua · Salesforce · custom REST

Delivery infrastructure →

MONEY FOUND

Your documents owe you money.

Leakage hides where nobody can query: duplicate payments, unbilled escalations, expiring windows. Once every field is typed and traceable, one pass over the corpus brings it back.

REGISTRY FINDINGSONE QUERY · WHOLE CORPUS
Invoice paid twice€18,400Same vendor, invoice number, and amount under two different scan filenames.
Price escalation never billed€31,200Index clause in the contract, unit price on the invoices: 14 months apart.
Auto-renewal caught in time€54,000Notice window read from the contract, flagged 96 days before it closed.
Credit note never applied€7,950Open credit with no matching deduction on any later invoice.
Freight billed above rate card€12,300Carrier invoice line against the contracted lane rate.
Found in one corpus pass€123,850

Representative findings from AP and contract portfolios. Every line traces to the value, the source line, and the scan that produced it.

RETRIEVAL

RAG where accuracy matters.

Vector search over raw text retrieves passages that look relevant. The registry retrieves values that are checkable — typed, scored, and traceable to the line that produced them.

When the answer feeds an ERP entry, a filing, or a decision that gets audited, the right-looking paragraph is not a data source.

Embedding retrievalRegistry retrieval
ReturnsText chunksTyped fields
CertaintySimilarity scorePer-field confidence
SourceApproximate passageLine + region of the scan
On re-queryRe-embed, re-rankRead from the registry
Audit trailValue → line → scan → reasoning

DOCUMENT ONTOLOGY

529 document types. Zero templates.

From Schedule K-1 to Bill of Lading (Ocean), from Notarial Deeds to QC Inspection Forms — Talonic understands the structure of every enterprise document type out of the box.

Schedule K-1, Form 1099-MISC, Form W-8BEN, Form 1040, Form 990, and 48 more

Purchase Order, Commercial Invoice, Pro Forma Invoice, Goods Receipt Note, Three-Way Match Report, and 48 more

Notarial Deed, Master Service Agreement, Statement of Work, Non-Disclosure Agreement, License Agreement, and 48 more

Commercial Register Extract, Certificate of Good Standing, Certificate of Incorporation, Annual Return, Director's Report, and 48 more

Patient Intake Form, Medical History Questionnaire, Informed Consent Form, Discharge Summary, Operative Report, and 48 more

QC Inspection Form, Certificate of Analysis (CoA), Certificate of Conformance, First Article Inspection Report, Material Test Report, and 48 more

Insurance Application, Policy Declaration Page, Insurance Policy Wording, Endorsement / Rider, Certificate of Insurance, and 48 more

Property Deed, Title Report, Land Registry Extract, Survey Plan, Zoning Certificate, and 48 more

Employment Application, Offer Letter, Employment Contract, Background Check Report, I-9 Employment Eligibility, and 47 more

For AI Agents

Built to be the data layer your agent reaches for.

Point your agent at Talonic and it can read any document like a database: typed fields, per-cell confidence, and provenance. Discovery surfaces and an MCP server are ready out of the box.

Connect via MCP

npx -y @talonic/mcp@latest
Start free, get an API key →

DIN SPEC 91491

We co-authored the standard.

DIN — Deutsches Institut für Normung

Talonic co-authored DIN SPEC 91491 with Fraunhofer IIS, Humboldt-Innovation, GIIC, and the German standards body. Europe's first standard for AI-ready data at the schema layer.

As the EU AI Act extends into data-layer compliance, DIN SPEC 91491 defines what schema-layer data readiness looks like. Enterprises aligning with the standard need an implementation that was built alongside it.

The standard we helped write. The implementation we ship.

See what Talonic would do to your data.

Send a representative sample of contracts, scans, case files, or operational documents. We run them through Talonic and send back your extracted data: field-level coverage, confidence, provenance, and a concrete recommendation for making your corpus reusable, within five business days.

P.S. Six fields, committed once, queryable forever. Your documents work the same way.

Not ready to share a sample?

Document testresponse within 1 business day · your data back within 5

Response within 1 business day.

By submitting, you agree to our Privacy Policy. Your data is processed on EU-resident infrastructure and never shared with third parties.

Frequently asked questions

What is Talonic?

Talonic is the schema layer for enterprise data. It transforms unstructured documents into schema-validated structured data with per-cell provenance and full audit trail.

What types of documents can Talonic process?

Talonic supports over 25 file formats including PDFs, scans, Word documents, Excel spreadsheets, images, emails, and more, classified into a 529-type document ontology.

Is Talonic GDPR compliant?

Yes. Talonic is GDPR and HIPAA compliant, ISO 27001 and ISO 42001 aligned, and co-authored DIN SPEC 91491. All data is processed in EU-resident infrastructure in Germany West Central.

What is DIN SPEC 91491?

Europe's first standard for AI-ready data at the schema layer, published in 2025. Talonic co-authored the standard alongside Fraunhofer IIS, Humboldt-Innovation, and GIIC.

How does Talonic differ from OCR or document parsing tools?

OCR converts pixels to text. Talonic classifies documents against a 529-type ontology, extracts schema-validated fields with confidence scores, reconciles entities into cases, and delivers typed data to systems of record with full audit trail.

Can I test Talonic on my own documents?

Yes. Send a representative sample of your documents and Talonic runs them within five business days. You receive a field-level analysis of extractability, confidence distribution, recommended schema, and a production roadmap.