AI document processing: scans to 1C and XML

Scans in, verified data out

The agent reads invoices, acts, waybills and design documentation. It extracts fields, shows where each one came from, checks them against your rules and exports to your system. A person reviews anything uncertain.

Documents arrive as pictures, but the work needs data

  • An accountant retypes details from invoice scans into 1C and checks totals by eye.
  • A mistake in an act surfaces only when the counterparty sends it back unsigned.
  • The state review needs XML that matches a schema, yet you have a pile of PDFs, DOCX files and scans. Someone assembles it by hand, with errors every time.
  • Incoming paperwork keeps growing, and hiring more people for it isn’t the plan.

What changes after launch

  • 01Staff review only doubtful fields instead of retyping whole documents
  • 02Errors in requisites and totals get caught before posting, not after
  • 03Strict formats like XML pass schema validation before they go out

How the agent works

  1. 01

    Take in the document

    From email, a shared folder, e-document exchange or a web upload. PDF, DOCX, photos and scans.

  2. 02

    Recognize the text

    Scans go through OCR. Document structure — tables, requisites, signature marks — is preserved.

  3. 03

    Extract fields

    AI finds the fields you need and records the location on the page and a confidence score for each. A field without a source is rejected.

  4. 04

    Check against rules

    Tax IDs against your directory, totals against line items, dates against the contract. Rules are written in code, not left to the model’s judgment.

  5. 05

    A person settles doubts

    Low-confidence fields and broken rules go to an employee, who confirms, corrects or enters the value manually.

  6. 06

    Export

    To 1C, CRM, e-document exchange or a file in the required schema. XML and other strict formats are built by code, not by the model, and validated before sending.

Which scale is yours?

One task, three different projects. Pick the one that looks like you and press “I want this” — your request arrives tagged, and we send an estimate in 1–2 days.

Small business

One document stream, say incoming invoices, that eats hours of an accountant’s week.

What we build

  • One document type and a fixed set of fields
  • Web interface: upload, review, download
  • Required-field checks and simple rules
  • Export to Excel or CSV
Integrations
No integrations: a file goes in, a spreadsheet comes out
Where it runs
Cloud
Timeline
Orientation: about a month

Mid-size business

common start

Invoices, acts, waybills and contracts flow in from email and e-document exchange, and several people check them.

What we build

  • Several document types with automatic sorting on intake
  • Matching against 1C directories: counterparties, contracts, items
  • Review queue for staff with roles
  • Document created in 1C or CRM after confirmation
  • Report on which fields get corrected by hand most often
Integrations
1C, CRM, e-document exchange, email, shared folders
Where it runs
Cloud or your own server
Timeline
Orientation: 1–2 months and up

Enterprise

common start

Documents hold personal data and trade secrets, exports go into strict industry formats, and security has its own requirements.

What we build

  • Recognition and model inside a closed perimeter on your servers
  • Document access by department and legal entity
  • Audit log: who confirmed a field and what they changed
  • Export to industry formats with XSD validation
  • SLA and support when forms and schemas change
Integrations
1C, document management, e-document exchange, digital signature, internal APIs
Where it runs
Closed perimeter on your servers; GPU hardware quoted separately
Timeline
Orientation: 2–3 months and up

Pilot from RUB 300,000. Beyond that we price by task and hours.

Questions

What about bad scans?

Before we start, we take a sample of your real documents, including the worst ones, and see what gets recognized. A field read with low confidence never moves on without a person. If a scan is unreadable, we’ll say so plainly: it needs a rescan.

Can the model get a total wrong?

It can, which is why we never rely on the model alone. Every field is tied to a spot in the document, totals and requisites are checked by rules, and a person reviews anything in doubt. In the XML project for construction documentation, the model only suggests fields; the XML itself is assembled by code against the XSD schema.

How is this different from plain OCR?

OCR gives you text. You need fields: who the supplier is, the total including VAT, which contract the invoice belongs to. AI extracts them from the recognized text, then rules check the result.

Where are the documents stored?

Depends on the setup. In the cloud version, on the project server. If documents contain personal data or trade secrets, we deploy everything on your servers and connect the model so data never leaves the perimeter. The architecture is designed for 152-FZ.

How soon is the first version?

A working version for one document type is usually ready in about a month. Complex document sets take longer: stage one of the construction XML project was scoped at 45 days. Cost is estimated by task and hours: we send a staged estimate within 1–2 days after the brief.

A quote in 1–2 days

Name the process that eats time. You get a pilot range, not a 40-slide deck.

Or message us