BlogProduct & tech

    Why AI agents need sources before they act

    Maike Penz · CEO & Co-Founder · October 6, 2026 · 7 min read

    A chatbot gives an answer, and a person reads it. An AI agent works differently: it searches, evaluates, searches again and builds a final result from intermediate findings. If it goes wrong in step two, the error turns up in step five as a neatly worded fact. That is why a source citation is not a courtesy to the reader where agents are concerned. It is the only point at which the system’s work can be checked at all.

    What sets an AI agent apart from a chatbot

    The term “AI agent” is currently used for almost anything. Here it means something specific: a system that breaks a task into several steps and decides for itself what happens next. It formulates search queries, reads the hits, discards some of them, searches again in a targeted way and finally assembles a result.

    For company knowledge, that is real progress. Many questions on the shop floor cannot be answered with a single search hit, because the answer is spread across an inspection report, a work instruction and an old email.

    The price: between question and answer lie decisions nobody has seen. Which file did the agent use, which did it discard? With a chatbot, a person reads every answer. With an agent, they only read the end.

    How one wrong hit runs through the whole chain

    An example, deliberately kept simple: someone in production planning asks what protective equipment is needed when decanting a new cleaning agent, and whether the container may be stored next to the acids in the hazardous materials store.

    The agent first looks for the safety data sheet (SDS) and finds two versions. An older one sits in the purchasing folder, a newer one in the quality management directory. In the newer one, the supplier has changed the classification. If the agent takes the old version, it derives the protective equipment from it, then the storage recommendation, and perhaps even a draft workplace instruction for the substance. Every step is internally consistent. Only the first one is wrong.

    There is little contrived about this. The EU REACH Regulation requires suppliers to update safety data sheets as soon as new information on hazards becomes available (Article 31). Anyone familiar with network drives that have grown over the years knows that the old copies rarely all disappear afterwards. How easily data sheets get mixed up is shown by a test run on the website of a mid-sized petrochemicals company: for five products, the data sheets belonged to other products. This article does not cover which obligations apply in your company. It is not legal advice.

    The treacherous part: the final result doesn’t sound uncertain. Without a citation at every step, a well-reasoned wrong answer looks exactly like a right one.

    Why language models would rather guess than stay silent

    Why do language models sound confident even when they are wrong? The research service of the German Bundestag summarizes the research as follows: a model calculates the most probable next word and is therefore under pressure to deliver a complete answer, even when facts are missing. Training methods also often reward guessing rather than admitting uncertainty.

    The usual countermeasure is called retrieval-augmented generation (RAG): before answering, the model is given matching passages from your own documents. This improves factual accuracy considerably, but it does not eliminate the problem. Errors still occur when the search returns the wrong document or the model processes the retrieved text incorrectly. An agent doesn’t call the search once, but several times in a row. Each call is another opportunity for exactly this error.

    Source citations are not automatically correct either. In a study by the BBC and the European Broadcasting Union on news questions, 45 per cent of AI assistants’ answers contained at least one significant error. The most common cause was faulty sourcing.

    This leads to a sober rule: a source you can’t check with one click is a claim with a footnote. What this means for the time it takes to get an answer you can actually work with is something we calculated in our article on searching for information in companies.

    How to recognize a usable citation

    Whether you use hAiner, another system or are just comparing providers: you can check these points yourself on every single answer.

    • It leads to the original. A file name like “Manual.pdf” is not a citation but a search task. You need to be able to open the document directly from the answer and see which repository it is stored in.
    • It covers the whole answer. If an agent relies on three documents, all three belong in the answer, not just the first.
    • It names contradictions. If two versions with different information exist, that belongs in the answer. The system must not quietly pick one of them.
    • It distinguishes between document and statement. If information comes from a conversation, it should say who said it, when and in what role.
    • It is honestly absent when there is none. “I found nothing on this in your documents” is a good answer. Better still if the system adds who in the company might know.

    How to test this with your own documents in an hour is described in the checklist in our comparison ChatGPT or Copilot for company knowledge.

    One limit applies to every system: whether a document version is valid is rarely stated in the text. That is governed by the control of documented information in the quality management system. Where it is properly maintained, a system can build on it. Where it is not, the system should not guess the status.

    Experience-based knowledge needs a sender too

    Some of the most important answers are not in any document. That a particular pump has to be started up more slowly after a longer standstill than the manual says may be known only to the maintenance technician who has worked on that line for years.

    When an agent passes on knowledge like this, the same rule applies. The citation is then not a PDF but a statement with a sender: who, in what role, when and in what context. This has two practical consequences. The next person can judge how much weight to give the statement. And they know whom to ask, as long as that person is still with the company.

    How to capture this kind of knowledge in the first place, before it disappears with a retirement or a job change, is described on our topic page Securing experience-based knowledge.

    How hAiner handles this

    Search in hAiner works agentically itself. It combines full-text search for exact terms, a semantic search that also finds differently worded passages, and queries across a knowledge graph in which documents are linked to people, equipment and processes. The system decides step by step which route fits the question.

    That is precisely why we apply a citation requirement without exception: every answer names the sources it relies on, together with the path where the document is stored. The original can be opened directly from the answer. How hAiner accesses SharePoint and network drives without copying the documents is described in our article on SharePoint access.

    If contradictory information turns up, hAiner points it out instead of quietly choosing one version. Statements from knowledge interviews are kept as a separate source, with the person and the time of the conversation. hAiner presents the outcome of a conversation to the person for approval before it counts as knowledge.

    And if there is no citation? Then hAiner doesn’t construct an answer, but refers you to the responsible person in the company who might know. A knowledge gap thus becomes a concrete question to a concrete person.

    We name one limit explicitly. Which of two contradictory versions is technically valid is something hAiner cannot decide if nobody in the company has determined it. That remains a decision for people. hAiner makes visible that it is due.

    We are happy to show you what a citation in hAiner looks like in practice, and where its limits are, in an initial consultation.

    Frequently asked questions

    What is an AI agent in a company?
    An AI agent is a system that breaks a task into several steps and decides for itself which step comes next. In knowledge management this usually means it searches several sources, evaluates the hits, searches again in a targeted way and assembles an answer from them. Unlike with a simple chatbot, people usually don’t see the intermediate steps.
    Do source citations prevent an AI from hallucinating?
    No. Source citations don’t prevent errors, they make them checkable. Even methods that give the model passages from your own documents in advance (RAG) reduce hallucinations but do not rule them out. What matters is that every citation can be checked against the original with one click.
    What makes a good source citation in AI answers?
    A good citation can be opened directly, shows where the document is stored, names all documents used and points out contradictory versions. If the information comes from a conversation, it includes the person, their role and the time. If there is no citation, the system should say so and ideally refer you to someone who might know.
    Can an AI tell which version of a document is valid?
    Only if that information is maintained in the company, for example through document control in the quality management system. If it is missing, an AI can at most show that several versions with different information exist. Which of them is valid must be decided by people.

    Sources

    Sound like your situation?

    Let's talk for 30 minutes about your concrete case. No obligation, no pitch deck.