SEO MachineSoftware, growth, AI & sales, done for you

ServiceAI agents

Document AI: ask your company's documents a question, get an answer with its source

Our document AI service builds an assistant that answers questions from your own documents — contracts, conditions, procedures, product sheets, past exchanges — and shows the passage each answer comes from. When the documents do not contain the answer, it says so. Technically this is often called RAG (retrieval-augmented generation): the assistant first finds the relevant passages, then writes an answer based only on them.

What we do for you

  • An inventory and clean-up of the documents the assistant should use, and of those it must not
  • Preparation of the corpus: conversion, splitting into passages, removal of duplicates and outdated versions
  • An assistant that retrieves the relevant passages and answers from them, with the source shown
  • An explicit "not in the documents" answer rather than a guess
  • Access rules so people only get answers from documents they are allowed to see
  • A test set of real questions with expected answers, run before launch and after each update
  • Delivery where your team works: a web page, the team chat or WhatsApp

Who it is for

  • Fits: teams that spend time searching long documents for the same answers (conditions, procedures, product details)
  • Fits: brokers, training organisations, law firms, support teams and anyone with a large body of written rules
  • Fits: companies that want a CRM or internal assistant to know their history
  • Does not fit: decisions that must not rest on an AI answer without a qualified person checking it
  • Does not fit: document sets that are contradictory and that nobody can say which version is right — that needs sorting first

What we have built

We prepared a corpus of 1,226 documents to serve as the memory of a CRM assistant: conversations and records turned into passages an assistant can search. We also built a proof of concept that answers questions about insurance underwriting conditions from the insurer's own documents. Both are anonymised here, and both taught us that most of the work is in the documents, not in the AI.

How we keep answers grounded

Language models can produce confident text that is wrong. Anthropic, one of the main model providers, documents the techniques that reduce this: give the model explicit permission to say it does not know; for long documents, have it extract word-for-word quotes before answering; make answers auditable by citing a supporting quote for each claim and retracting any claim without one; and restrict it to the documents provided rather than its general knowledge. Anthropic adds that these techniques reduce hallucinations but do not eliminate them, and that critical information should always be validated.

We build on exactly those principles: answers come with the passage they rely on, "not found" is an acceptable answer, and for high-stakes uses a person checks before acting.

How a project runs

  1. Questions first. We collect the questions your team really asks. They become the test set.
  2. Documents next. We list the sources, remove outdated versions and decide what is in and out of scope.
  3. Prepare. Conversion, splitting into passages, metadata (date, owner, who may see it).
  4. Build and test. The assistant answers the test set; we compare with expected answers and fix retrieval or instructions until results are acceptable to you.
  5. Deploy. Where your team works, with access rules.
  6. Maintain. New documents are added; the test set is run again after each change.

Personal data and confidentiality

Company documents often contain personal data: client files, staff records, emails. The CNIL's guidance for AI systems asks for a well-defined purpose, data that is "adequate, relevant and limited to what is necessary", clear information to the people concerned and a retention period set in advance; the GDPR adds that data must not be kept longer than necessary (Article 5). In practice we leave out documents the purpose does not need, respect existing access rights, and tell you which AI provider processes the text and under which terms.

Search box, chatbot or document assistant?

ToolWhat it doesLimit
Search in your driveFinds files containing wordsYou still read the whole file
Generic AI chatbotWrites fluent answersNot tied to your documents; can invent
Document assistantFinds the passages and answers from them, with the sourceOnly as good and current as your documents

Questions we get

Will the assistant make things up?

It is built to answer only from your documents, to show the passage it used and to say when the answer is not there. That reduces errors but does not remove them, so important answers should be checked by a person.

Which documents can it use?

Most text documents: PDFs, Word files, web pages, exported emails or CRM notes. Scanned pages need text recognition first; the inventory tells you whether yours are usable.

Can different people see different answers?

Yes. We carry over access rules, so someone only gets answers from documents they are allowed to read.

How do you know it works?

With a test set of real questions and expected answers, agreed with you before launch and run again after every update.

Where will my team use it?

Where they already work: a web page, the team chat or WhatsApp.

Sources

  1. Anthropic — Reduce hallucinations (checked 2026-10-06)
  2. CNIL — AI: how to comply with the GDPR (checked 2026-10-06)
  3. GDPR, Chapter II (Article 5) — CNIL edition (checked 2026-10-06)

Want to know what we would do first?

Tell us your business, your town and your website. We come back by email with a first plan: growth, AI agents, sales or all three.

Get your growth plan
Get your growth plan