On-premise LLM deployment: a private ChatGPT for business
A corporate ChatGPT that keeps data inside
We run an open language model on your servers, connect it to your documents and systems, and give staff a chat and an API. Answers cite their source, access is role-based, data stays inside the perimeter.
What we build
Public chatbots are convenient, but not for contracts and customer data. We deploy an open model (Llama, Qwen, Mistral and other open-weight models that handle your languages) on your servers and add search over your documents - RAG. An employee asks a question, the model finds the right clause in a policy or contract and answers with a link to it. We pick the model for the task and the hardware, and test answer quality on your real questions before launch.
Corporate AI chat
A chat for staff over policies, contracts and manuals. Answers cite the document and nothing goes to an outside provider.
RAG over your documents
Document indexing, knowledge base updates and access rights: each employee sees only what they are allowed to.
API for your systems
The same model serves your CRM, 1C and internal services: email parsing, request classification, document drafts - inside the perimeter.
Private AI agents
The AI agent scenarios - sales, support, routine - with models and data kept on your servers.
What the work includes
Priced per task, estimate within 1–2 days after the brief. GPUs and servers are listed separately.
- Model choice for the task, language and hardware - compared on your questions
- Inference deployment on your servers or in your private cloud
- Knowledge base: document upload, indexing and updates
- Chat for staff and an API for systems
- Roles via SSO and a query audit log
- A 4–8 week pilot with acceptance criteria
- Documentation and handover to your team
How we work
- 01
Requirements
We clarify which data must not leave: personal data, trade secrets, critical infrastructure. We record your security team’s requirements and the use case behind the project.
- 02
Perimeter architecture
Where models, the knowledge base and logs live, who can access what, and how the perimeter connects to your systems. The design is agreed with security before any hardware is bought.
- 03
Pilot inside the perimeter
4–8 weeks on a real task: model, documents, roles. Acceptance criteria - share of correct answers with a source, speed, hand-offs to people - are agreed upfront.
- 04
Fine-tuning and scaling
If a base model with document search is not enough, we fine-tune it on your data. Users, use cases and capacity grow with the load.
- 05
Support
Monitoring, model and knowledge base updates, incident handling. The perimeter and the code stay yours.
Questions
Which models do you use?
Open-weight models you can install on your own servers: Llama, Qwen, Mistral and others. We choose by quality on your questions and by the hardware available.
Will a private model be worse than ChatGPT?
On general knowledge, sometimes. On questions about your documents, search over the knowledge base and clean data decide: there a private model with RAG answers precisely and cites the source.
What do you measure in the pilot?
Share of answers with a correct source, hand-offs to people, response time and the pilot group’s rating. Criteria are agreed before the start.
Can we start without our own GPUs?
Yes: the pilot runs on available hardware or a rented server, and the purchase is planned from the real load.
The perimeter is live - next comes support
After launch we support the perimeter: model and knowledge base updates, monitoring, incidents and improvements under a contract or an SLA.
Other development areas
- Closed AI perimeterA perimeter design for models and data: roles, audit logs, network segments. Built around data protection law and your security rules - before any GPU purchase.
- Model fine-tuningFine-tuning an open LLM on your data: terminology, answer format, classification. Inside your perimeter, with a quality check.
- GPU and infrastructureGPU server sizing for your model and load, site readiness and orchestration. Hardware for the task, not the top of the catalogue.