One questionnaire per working day, and everything depends on a single approver.
A German specialty chemicals producer supplies materials into very different applications. Each one brings its own regulatory questions, and customers ask them in writing. Around 250 questionnaires arrive per year, about one per working day. The regulatory team that answers them is small.
Nothing is standardized at the input: free text in an email, a Word form, an Excel sheet, a PDF, and increasingly a customer portal that wants to be filled in field by field. Every answer is assembled by hand from three separate sources: the company’s own statements and safety data sheets, old customer cases, and the original legal text.
That costs about half an hour of pure handling per case on average, up to a full working day in extreme cases. A good week passes before the customer has the answer. The bigger pain, though, is the concentration risk: every case goes to a senior approver who is heading towards retirement, and his judgment is written down nowhere.
Collect the real material first, then make the expert’s judgment explicit.
The team had already tried once to write down which answers may go out unchecked. The rulebook grew so broad that they abandoned it. Exactly that became the condition: no rigid rulebook, every case still through the senior approver, and the question about formulation data openly on the table before anyone writes code.
- Question catalogue before the first call: We first clarified which questions actually arrive, in which formats, and from which sources they are answered today.
- Discovery call on a pre-drawn workflow: We came into the meeting with the as-is workflow already drawn. The team corrected the drawing instead of describing a process from scratch. Corrections are more precise than descriptions.
- Real material instead of examples: Two things mattered most: the request log of a full year with categories and minutes recorded, and five real questionnaires selected by archetype, from a multi-round email thread to an absence declaration covering about 100 substances.
- Deliberately broad on sources: All three knowledge sources from day one. An AI employee that reaches only one of them merely shifts the search work instead of taking it away.
The AI employee makes the proposal. The human decides at two points.
Intake runs through the team’s shared mailbox, no new mailbox and no new habit. The AI employee reads the questionnaire in any format, isolates the individual questions, and matches each one to the product concerned. For every question it creates an answer proposal from the three knowledge sources and attaches a source citation, a short rationale, and a confidence level. Matching attachments such as safety data sheets are referenced automatically.
The confidence model is the heart of the build. High means a direct in-house source answers the question. Medium means an earlier answer plus a similar source support it. Low means research or reasoning is needed. Critical means a human has to decide. Two topics are hard-wired as critical, whatever the AI employee finds: food contact and automotive dossiers. The clerks review the flagged points, then the senior approver releases every case. Every released answer flows back into the knowledge base together with its question and its source.
Customer questionnaire on one product line, 7 pages, 11 fields, including a PFAS absence declaration covering about 100 listed substances. Received as a PDF attached to an email thread.
One field is flagged for review. Food-contact questions always go to a human, no matter how confident the AI employee is. On top of that comes the recurring structural problem: customer forms often demand a yes or no where the honest answer carries conditions. The AI employee may propose the wording, only a human may release it.
That was before, this is today.
The activity changed, not just the duration: before, people searched; today, they review. Conservatively calculated, that saves more than half of the handling time per questionnaire.
| Before | Today | |
|---|---|---|
| Research | By hand across three separate sources | One answer proposal per question, with source citation |
| Time per questionnaire | About 30 minutes on average, up to 8 hours in extreme cases | 10 to 15 minutes of reviewing instead of searching |
| Turnaround | A good week | A few days |
| Release | Without documented criteria | Confidence level per answer, the senior still releases every case |
| Knowledge | In the head of one person close to retirement | Released question, source, and answer stored in the knowledge base |
| Legal status | Checked ad hoc | Newsletter monitoring flags the affected statements |
* Baseline taken from the team’s own request log across a full year, around 250 cases with categories and minutes recorded. Savings conservatively calculated: about 30 minutes of searching and writing per questionnaire before, 10 to 15 minutes of reviewing today. That is 15 to 20 minutes saved per case, so 62 to 83 hours per year at 250 questionnaires. The extreme cases of up to 8 hours are not included in this calculation.
How do you build an AI employee like this, technically?
The knowledge from this project to take away, whether you build with us or on your own:
Reading: the questionnaire is a source of questions, not a form
Language models like Claude from Anthropic or GPT from OpenAI read email free text, Word forms, Excel sheets, PDFs, and portal exports alike, and pull out the individual questions. The container format is deliberately not part of the logic. Intake connects through the Microsoft Graph API to the existing shared mailbox in Outlook, and the statements stay in SharePoint.
Answering: three sources, every answer with a citation
A proposal is built from your own statements and safety data sheets, from previously answered customer cases, and from the original legal text. Every answer carries its location in the source and a short rationale. Without the citation the proposal would be worthless, because an approver only reviews faster than he would write himself when the evidence is right there.
A confidence model instead of an automation rate
A rate hides exactly what a compliance team needs to see. Instead of deciding upfront which cases are safe, the AI employee states per answer how well it is backed: high for a direct in-house source, medium for an earlier answer plus a similar source, low when research is needed, critical when a human has to decide. Uncertainty becomes visible information instead of a hidden percentage.
Staying current: monitoring the legal status
A dedicated mailbox subscribes to the relevant regulatory newsletters. The AI employee reads announced changes, for example a new SVHC candidate entry, a tightened PFAS restriction, or a changed RoHS exemption, and flags every in-house statement affected by it. REACH, TSCA, RoHS, and PFAS each bring their own answer patterns, so monitoring runs per framework.
UX: the interface decides adoption
The review screen shows the question from the customer form on the left and the proposed answer with its source citation on the right. The release flow has two stages: first the clerk on the flagged points, then the senior on the whole case. And every release writes the question, the source, and the answer back into the knowledge base, which gets denser with every case.

What does it cost in comparison?
Superkind charges per use case. The price grows with request volume, not with headcount. Here is the honest comparison:
| Regulatory specialist | External consultants | Superkind AI employee | |
|---|---|---|---|
| Cost | 60,000 to 85,000 € per year | Day rates in the range of 1,200 to 2,000 € | Price per use case, a fraction of a full-time position |
| What is included | The whole case, by hand | Single cases and expert opinions | Splitting questions, proposing answers with sources, referencing attachments |
| Scales with | More staff | More days booked | Request volume, without new positions |
| Where the knowledge stays | In one person’s head | With the external firm | In your own knowledge base |
| Exceptions | Human does everything | Commissioned separately | Flagged and sent to a human |
| Rollout | Recruiting and onboarding | Briefing per case | 2 to 3 weeks to the first productive version |
The honest comparison is the full cost of today’s routine: the search time per questionnaire, the week your customers wait, and the risk of holding regulatory judgment in a single head.
What we learned from this project.
The most valuable dataset was an unremarkable Excel file. The team had logged its requests for a full year, with category, product, and minutes spent. That table told us more than any workshop: which question types actually dominate, where the extreme cases sit, and what a realistic baseline looks like. We now ask for it early in every project.
The second lesson changed how we scope projects like this. Writing down the rulebook had already failed here, because it broke under its own breadth. A confidence model works exactly where a rulebook fails, because it allows uncertainty. The AI employee does not need to know whether a case is safe. It only needs to say how well its answer is backed. Automation that is allowed to admit uncertainty gets accepted by specialists.
What it is not suited for: If only a handful of questionnaires arrive per year, the routine is not your bottleneck. Without current and findable statements the first project is a document project, not an AI project. And in areas where every answer demands a fresh expert opinion, the AI employee can prepare but cannot shorten the work.

