What about bad scans?
Before we start, we take a sample of your real documents, including the worst ones, and see what gets recognized. A field read with low confidence never moves on without a person. If a scan is unreadable, we’ll say so plainly: it needs a rescan.
Can the model get a total wrong?
It can, which is why we never rely on the model alone. Every field is tied to a spot in the document, totals and requisites are checked by rules, and a person reviews anything in doubt. In the XML project for construction documentation, the model only suggests fields; the XML itself is assembled by code against the XSD schema.
How is this different from plain OCR?
OCR gives you text. You need fields: who the supplier is, the total including VAT, which contract the invoice belongs to. AI extracts them from the recognized text, then rules check the result.
Where are the documents stored?
Depends on the setup. In the cloud version, on the project server. If documents contain personal data or trade secrets, we deploy everything on your servers and connect the model so data never leaves the perimeter. The architecture is designed for 152-FZ.
How soon is the first version?
A working version for one document type is usually ready in about a month. Complex document sets take longer: stage one of the construction XML project was scoped at 45 days. Cost is estimated by task and hours: we send a staged estimate within 1–2 days after the brief.