Wrong data sheets on your website: found with a single question
In a test run for a mid-sized supplier of petrochemical specialty products, we gave hAiner nothing but the company's public website. No internal systems connected, no project set up, just the pages any visitor can see. Then a single, deliberately general question: "Where are data sheets missing?" The answer was a list. And that list was revealing.
One question, five hits
The product range on the website was well maintained: every product had its own page, every page had a data sheet attached. Checking the list still turned up five products where the attached data sheet belonged to a different product. Formally flawless PDFs: they open, they carry the company logo and layout, they are genuine data sheets. They are simply linked under the wrong product.
There is a simple reason why such a general question is what surfaces these cases. To a system that reads content, another product's data sheet is not a data sheet for this product. From that perspective, a misfiled document is simply a missing one.
These documents were not sitting on some internal drive. They were public; any customer could have been working from them.
Why nobody notices
A broken link gets noticed, because it produces an error. A wrong PDF produces none: it opens, it looks professional, it is a real data sheet. To catch the mistake, someone would have to read the content of the document against the product it is attached to. In day-to-day work, nobody does that. Website maintenance checks availability, not belonging.
There is a second layer. Knowledge about how documents are filed is itself experiential knowledge. Why there are two near-identically named product variants, which file is the authoritative one, which page is still left over from the old shop system: none of that is written down anywhere. The colleague who built the pages knows it. When she leaves, the map leaves with her.
Above all, day-to-day work has no reader who reads every document at once and could notice contradictions. People read selectively. To them, this kind of error is invisible.
What wrong data sheets cost
In the chemical industry this is not a matter of tidiness. If a customer designs a process around the values in the wrong technical data sheet, the end of that story is a complaint, and the question of whose document they trusted. With safety data sheets it gets sharper: REACH obliges the supplier to provide the safety data sheet for the substance or mixture in question (Art. 31; in Germany specified further by TRGS 220). A formally valid but misfiled SDS means, in practice, that the customer has no correct safety data sheet for this product at all.
Quality management has caught up with the topic as well. Since the 2015 revision, ISO 9001 explicitly requires under "organizational knowledge" (clause 7.1.6) that companies determine, maintain and protect the knowledge their processes need. Misfiled product documents are the precise opposite, which makes this an audit topic too.
The quieter damage is internal. A sales team that no longer trusts its own website goes back to calling that one experienced colleague. The single point of knowledge you set out to dissolve keeps growing instead.
What you can do, even without AI
- Name one owner per document set. Safety data sheets and technical data sheets each need an owner, not "the team".
- Define one authoritative source. The website pulls documents from one system; every manually uploaded copy is a future error.
- Unambiguous file names tied to the product. A datasheet_final_v2.pdf is a mix-up waiting for its moment.
- Spot checks mean reading content, not clicking links. Open ten product pages: does the PDF carry the same product name as the page? Are the values plausible for it?
- Hand over the document map. When long-serving employees leave, the handover has to include knowing where things are and which version is authoritative.
If you want to tackle this at a fundamental level: ISO 30401 is now a dedicated international standard for knowledge management systems.
The first reader that sees everything
What is remarkable about this case is how little it took: no integrated system, no interfaces, no training. Just the public website as the data basis and one generic question. A link checker would have found nothing, because it only verifies that a document is reachable. A CMS only sees that a field is filled. Whether the content matches the product is something only a reader can see who understands both and has every page in view at the same time.
How to dissolve single points of knowledge like these systematically is set out on the topic page securing experiential knowledge.
That is exactly why this test works on any website, including yours: the knowledge-loss check takes five minutes, or you can arrange an initial conversation directly. hAiner is hosted entirely in Germany, in line with the GDPR.
Frequently asked questions
- How do I find misfiled data sheets on our website?
- In the short term, by spot check: open product pages and read the PDF content against the product, meaning product name, values, variants. Doing it systematically needs a system that reads and compares content. The case described here shows that the public website alone is enough of a data basis for that.
- Why doesn't a link checker find errors like this?
- Because it only verifies that the target is reachable. It does not check what is inside. The mix-up is correct on every technical level: valid link, valid PDF, genuine data sheet. The only thing that is wrong is which content is assigned to which product, and that can only be checked by reading.
- Do we have to roll out a system first?
- No. The test run used publicly accessible content only, with no internal systems connected. That makes this kind of analysis the lowest-threshold entry point: it produces meaningful findings before any internal project starts.
- Is this a process problem or an AI problem?
- Both. Owners, an authoritative source and a review cycle are process work and always worth doing. AI replaces the part that process cannot solve: the single reader who reads everything at once. (Note: this article is not legal advice.)
Sources
Sound like your situation?
Let's talk for 30 minutes about your concrete case. No obligation, no pitch deck.