Nick’s AdvisoryHR Consulting

Nick’s Advisory

AI & HR

AI-Generated Employment References in Germany: What Is Allowed

September 10, 2026 · 15 min read · by Nick, former Head of People

Nick, former Head of People and founder of Nick’s Advisory

Few HR tasks look as suited to AI at first glance as the German employment reference letter. The Arbeitszeugnis is a formalised document with a fixed structure, a coded language that has hardened over decades, and a pile that never shrinks during a restructuring. So the question from HR teams has become routine: can we simply have a language model write these?

The short answer is: the draft yes, the assessment no. And between those two sentences lies exactly the area where HR teams are currently making mistakes, because the question is decided by employment law and data protection law rather than by technology. The reference letter remains a statement by the employer about a person, with evidentiary weight, liability consequences and a statutory frame around it.

This article sorts out what is permitted when AI enters the reference letter process, what triggers works council co-determination, what may go into a prompt and what may not, and what a workable internal rule looks like. It does not replace legal advice in an individual case: the specific design belongs with your data protection officer, your works council and your employment law advisers.

The short answer: draft yes, assessment no

AI may take over the draft in the reference letter process, propose the structure, harmonise phrasing and assemble text blocks from existing material. What it may not take over is the assessment itself: the decision about how performance and conduct are rated has to be made, owned and signed by a human being. The reason is not caution. It is the legal nature of the document.

The Arbeitszeugnis is neither an administrative act nor an automatic output. It is a statement by the employer about a specific person that, in a dispute, is examined by a labour court for truth and completeness. Whoever did not form that judgement cannot substantiate it either. That is where AI-supported reference processes fail when they go too far: not on language, but on the burden of proof.

In practice this means a clear division of labour. The human sets the rating before the AI sees anything. The AI writes along that specification. The human checks, corrects and signs. Reverse the order and read the rating out of the generated text, and you have handed responsibility to a tool that cannot carry it.

The division of labour in three sentences

  • Rating first, made by a human The grade for performance and conduct is fixed before anyone writes a prompt.
  • Wording by the AI Structure, language, harmonisation, completeness check against the list of duties.
  • Review and signature by a human Owned by a person who can actually assess the performance described.

What German law prescribes for the reference letter

The frame sits in the Trade Regulation Act, and it is shorter than most guidebooks on the subject. On termination of an employment relationship, employees are entitled to a written reference containing at least details of the type and duration of the work performed. On request it extends to performance and conduct during the employment. That is the difference between a simple and a qualified Arbeitszeugnis.

The sentence that matters most for AI sits in paragraph 2: the reference must be worded clearly and comprehensibly, and it must not contain features or formulations whose purpose is to make a statement about the employee other than the one apparent from its external form or wording. That prohibition on coded signals applies to generated text exactly as it applies to text you wrote yourself.

Since the Fourth Bureaucracy Relief Act came into force on 1 January 2025, there is also a change many processes have not yet reflected: the reference may be issued in electronic form with the employee's consent. Before that, electronic form was expressly excluded. The electronic route requires a qualified electronic signature, and without consent it stays on paper.

On top of that sits the case law that sets the assessment benchmark. The Federal Labour Court has held that the formulation stating the work was performed “to our full satisfaction” constitutes an average rating corresponding to the middle grade on the satisfaction scale. Anyone demanding an above-average rating carries the burden of pleading and proof for it, and an employer issuing a below-average rating must substantiate that.

Two further points from the same line of case law: there is no entitlement to a closing formula with thanks and good wishes; where the employee objects to one that has been included, the only claim available is for a reference without that formula. And the signatory must be someone who, from a third party's perspective, is suited to take responsibility for the assessment, which in practice means a more senior manager with authority to give instructions.

Sources: Section 109 Trade Regulation Act: reference letters · Haufe: Bureaucracy Relief Act IV, employment law changes as of 1 January 2025 (German) · Federal Labour Court, judgment of 18 November 2014, 9 AZR 584/13 (burden of proof on the grade) · Federal Labour Court, judgment of 11 December 2012, 9 AZR 227/11 (no claim to a closing formula)

Where AI helps in this process, and where it hurts

The biggest gain is not where most people look for it. It is not faster writing, it is consistency. References from one company often differ more by the person who wrote them than by the performance they describe. That is a fairness problem and, in case of doubt, a discrimination risk, because differences in assessment then attach to the manager rather than to the work.

The second real gain is the completeness check. A model that has the role's list of duties and the agreed grading in front of it reliably finds what is missing: unmentioned leadership responsibility, an omitted project phase, a conduct assessment with no reference to colleagues and superiors. Those gaps are the most common reason for demands to amend a reference.

The greatest damage arises in two places. First, when the model sets the grade: language models reproduce the coded reference language they learned and pick superlatives by textual pattern rather than by the facts on file. An automatically produced “always to our fullest satisfaction” is not an assessment, it is a linguistic reflex, and it still binds you.

Second, when generated formulations are adopted unchecked that carry a different meaning inside the reference code than intended. The law prohibits exactly such features and formulations. Adopting rather than authoring them changes nothing about your liability, and in a dispute the explanation that the AI suggested it is worth nothing.

Risky use

Bullet points and personnel data into an open tool, the grade emerges from the generated text, phrasing is adopted because it sounds professional, signed by HR without first-hand knowledge of the performance.

Defensible use

Rating documented by the manager, pseudonymised input into a contractually secured system, AI drafts and checks completeness, a more senior manager reviews the substance and signs.

The difference is not the tool. It is the point in the process where a human decides.

Sources: Section 109(2) Trade Regulation Act: clarity and prohibition of hidden statements · Federal Labour Court, judgment of 4 October 2005, 9 AZR 507/04, seniority of the signatory (dejure.org)

Data protection: what may go into the prompt

A draft reference contains personal data by definition, and specifically employee data. Processing it is permissible where necessary for carrying out or ending the employment relationship. German law expressly counts applicants and people whose employment has ended as employees for this purpose. Producing a reference that is legally owed falls inside that frame. Uploading an entire personnel file into an arbitrary tool does not.

The first decision is therefore not the prompt, it is the system. A tool whose inputs are used for training, or which runs outside a data processing agreement, is unsuitable for reference letters no matter how well it writes. Clarify before the first draft where data is processed, how long it is stored and who can access it.

The second decision is data minimisation in the individual case. For a usable draft, a model needs the role, the period, the duties, the projects and the agreed rating. It needs neither the name nor the date of birth, neither sickness records nor details of parental leave, disability status, pay or religious affiliation. Those belong in the finished document or in the file, not in the prompt.

The third decision concerns automation. The General Data Protection Regulation gives people the right not to be subject to a decision based solely on automated processing which produces legal effects concerning them or similarly significantly affects them. Where exceptions apply, suitable measures are required, at minimum the right to obtain human intervention, to express one's point of view and to contest the decision.

Whether a reference letter crosses that threshold depends on the case and is not settled law. In practice the question is easy to defuse: if a human sets the rating and reviews the substance of the text before signing, there is no solely automated decision. The German data protection authorities' guidance on artificial intelligence formulates the same idea as a process requirement: purpose limitation, data minimisation, transparency towards the people concerned, and a human final decision on significant matters.

  1. 1

    Settle the system, not the prompt

    Data processing agreement, storage location, retention period, no training on your inputs. Without those four, no reference drafting.

  2. 2

    Reduce the input

    Role, period, duties, projects, agreed rating. No name, no date of birth, no health or social data.

  3. 3

    Specify the rating instead of asking for it

    The grade goes into the prompt, it does not come out of it. This is the single most important rule.

  4. 4

    Document the human review

    Who reviewed, when, and what changed. In a dispute this is your evidence that no automated decision took place.

  5. 5

    Create transparency

    State in the employee privacy notice that and for what purpose AI support is used.

Sources: Section 26 Federal Data Protection Act: processing for employment purposes · Article 22 GDPR: automated individual decision-making · German Data Protection Conference: guidance on artificial intelligence and data protection (German)

The AI Act: is a reference assistant a high-risk system?

The European AI Act expressly classifies AI systems in the employment context as high-risk. Annex III point 4 names two categories: first, systems intended for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications and to evaluate candidates; second, systems intended to make decisions affecting terms of work-related relationships, promotion or termination, to allocate tasks, or to monitor and evaluate the performance and behaviour of persons in such relationships.

The second category is the relevant one for reference letters. What matters is the intended purpose: a system that evaluates performance and behaviour falls inside it. A text tool that writes out an assessment a human has already made does not pursue that purpose. This is exactly why the order of steps in your process is not a stylistic question but the switch between two very different regulatory levels.

On timing, something has shifted that many compliance plans have not caught up with. Obligations for stand-alone high-risk systems under Annex III were originally to apply from 2 August 2026. With the digital omnibus package, which entered into force on 27 July 2026, that date moved to 2 December 2027, and for systems embedded in regulated products under Annex I to 2 August 2028. The obligations themselves remain; only the starting point moved.

One obligation, however, has been in force for a while and is still widely overlooked: providers and deployers of AI systems must take measures to ensure a sufficient level of AI literacy among their staff. That rule has applied since 2 February 2025 and reaches every company using AI in daily work, regardless of risk class. For HR it means that drafting references with AI requires a documented briefing, not a verbal recommendation.

For a sense of scale on penalties: infringements of the prohibited practices can draw up to 35 million euro or 7 per cent of worldwide annual turnover, breaches of provider and deployer obligations up to 15 million euro or 3 per cent, and incorrect or incomplete information up to 7.5 million euro or 1 per cent. For small and medium-sized enterprises, the lower of the two figures applies in each case.

The timetable moved, the obligations did not. Setting the process up cleanly now means not rebuilding it in 2027.

Sources: AI Act, Annex III: high-risk areas, point 4 on employment · AI Act, Article 4: AI literacy · AI Act, Article 99: penalties · AI Act implementation timeline · Digital Omnibus on AI: amendments to the AI Act, in force since 27 July 2026

Works council: when co-determination applies

The most common mistake when starting with AI-supported HR processes is informing the works council once the tool is already running. The Works Constitution Act provides several points of attachment, and three of them were expressly extended to cover artificial intelligence.

The employer must inform the works council about the planning of work procedures and workflows including the use of artificial intelligence, and must discuss intended measures with it early enough for the council's proposals and concerns to be taken into account in the planning. Early enough here means before the decision on a tool, not before the rollout.

Where artificial intelligence is used in drawing up selection guidelines for hiring, transfers, regrading and dismissals, the co-determination rights on selection guidelines expressly apply to that as well. And where the works council has to assess the introduction or use of artificial intelligence in order to perform its duties, bringing in an external expert is deemed necessary. That clarification removes the usual argument about who pays for the expert.

Particularly relevant for reference letters is the classic provision on monitoring technology: the works council has a co-determination right on the introduction and use of technical devices designed to monitor the behaviour or performance of employees. A pure drafting tool generally does not fall under it. A system that pulls performance data from other sources in order to derive assessments certainly does.

The pragmatic route, from a head-of-people perspective: settle AI use in HR once in a works agreement instead of negotiating again with every new tool. The agreement describes permitted use cases, prohibited inputs, the human final decision, retention, and the involvement process for new tools. It costs a few weeks once and then removes friction permanently.

Four points of attachment in the Works Constitution Act

  • Information and consultation on planning Work procedures and workflows including the use of artificial intelligence, in time to shape the decision.
  • Co-determination on selection guidelines Applies expressly where artificial intelligence is used in drawing them up.
  • External expert for the works council Deemed necessary where the council has to assess the introduction or use of AI.
  • Co-determination on monitoring technology Devices designed to monitor behaviour or performance require the council's agreement.

Sources: Section 90 Works Constitution Act: information and consultation, including AI · Section 95 Works Constitution Act: selection guidelines, paragraph 2a on AI · Section 80 Works Constitution Act: experts, paragraph 3 on AI · Section 87 Works Constitution Act: co-determination, paragraph 1 number 6

The burden of proof stays with you

When a reference letter ends up in front of a labour court, the argument is almost never about grammar. It is about the grade. And the allocation there bears directly on AI use: whoever demands an above-average rating must plead and prove the facts supporting it. An employer issuing a below-average rating must substantiate it. Where neither can be established, the court lands on an average rating.

The anchor point for that is fixed in language: the formulation stating that the work was performed “to our full satisfaction” corresponds to the middle grade of the satisfaction scale, roughly a C. Everything above it needs justification, everything below it even more so. A language model knows this scale as a textual pattern and picks superlatives by frequency, not by what is on file.

From this follows the only rule that truly matters in the process: the grade is set before the prompt, documented, and the prompt states it. Take it from the generated text instead, and you may have issued an assessment for which no basis exists in the personnel file, and you will still have to defend it.

A second point concerns equal treatment. If generated references deviate systematically across groups of people, for instance because the model reacts differently to part-time details, parental leave or names, a discrimination risk arises. And the burden of proof there is settled: where a person proves indications suggesting less favourable treatment on a protected ground, the other side must prove that no breach occurred. A quarterly sample check across your generated references is therefore not a luxury but risk provisioning.

Below averageExcellent
The highlighted step is the legal zero point. Everything above and below has to be substantiated by someone, and that someone is not a model.

Sources: Federal Labour Court, judgment of 18 November 2014, 9 AZR 584/13 · Section 22 General Equal Treatment Act: burden of proof · Section 109 Trade Regulation Act: reference letters

An internal rule that fits on one page

Most AI policies in HR departments fail because of their length. Nobody who is under time pressure to write a reference reads twelve pages of definitions. What works is a rule that fits on one page and sits where the work happens, which is inside the reference template itself.

The rule has to answer four questions: which system may be used, what may go into it, who decides what, and what gets documented. Everything else is commentary. Once those four are answered, the process is in a state you can explain, both under data protection law and under employment law.

It also matters not to write the rule as a list of prohibitions. A policy that only says what is forbidden produces shadow usage in private accounts, and with it exactly the risk it was meant to prevent. So state explicitly what is allowed, and provide a tool in which it is allowed.

  1. 1

    1. Approved system

    Only the named, contractually secured tool. Private accounts and open tools are excluded for reference letters.

  2. 2

    2. Permitted inputs

    Role, period, duties, projects, agreed rating. No names, no health data, no social data, no pay information.

  3. 3

    3. Rating set by the manager

    The grade for performance and conduct is fixed and documented before the tool is used. It is never taken from the generated text.

  4. 4

    4. Human review

    Completeness, truthfulness, no hidden statements. Reviewed by someone who can actually assess the performance.

  5. 5

    5. Signature with standing

    A more senior manager with authority to instruct signs. Electronic only with consent and a qualified signature.

  6. 6

    6. Documentation

    Who reviewed, when, what changed. One field in the template is enough.

  7. 7

    7. Quarterly sample

    Check ten references against the duty list and the rating, looking for systematic deviations between groups.

  8. 8

    8. Involvement settled

    Works council informed and consulted, data protection officer involved, employee privacy notice updated.

What HR actually gains from this

The time saved on an individual reference is real, but it is not the actual return. Setting the process up this way mainly buys consistency: references from different departments become comparable because they follow the same structure and use the same rating anchors. That reduces demands for amendment, and it reduces the likelihood that a difference in assessment is later read as discrimination.

The second return sits in the separation phase itself. A reference letter that is ready early and properly written is one of the few things that genuinely improves a separation for the person concerned. Leaving it lying saves nothing and risks turning a closed matter into an open account.

And the third return is the least spectacular one: a cleanly regulated AI use case in one clearly bounded place is the best preparation for all the others. Once you have settled which system is approved, what may go into it, who decides and what is documented, you can carry those four answers into the next use case instead of starting from scratch every time.

One closing note that applies to the whole article: it describes the legal frame in general terms and does not replace legal advice in an individual case. The specific design of an AI-supported reference process belongs with your data protection officer, your works council and your employment law advisers before the first reference is produced this way.

Sort out AI use in HR

A free initial conversation covering your HR processes: where AI already carries weight, where it becomes legally risky, and which rule will actually be followed in your team.

Sort out AI use in HR

Frequently Asked Questions

Can a German Arbeitszeugnis be written with AI?

The draft yes, the assessment no. AI may structure, phrase and check completeness. The rating of performance and conduct must be set, substantively reviewed and signed by a human, because the employer has to substantiate it in a dispute. The specific design belongs with your data protection officer, works council and employment law advisers.

What data may go into the prompt?

Only what the draft requires: role, period, duties, projects and the previously agreed rating. Name, date of birth, health and social data, pay, parental leave or disability status do not belong there. A precondition is a system covered by a data processing agreement in which your inputs are not used for training.

Is an AI reference assistant a high-risk system under the AI Act?

That depends on the intended purpose. Annex III point 4 covers systems used for decisions affecting work-related relationships or for evaluating performance and behaviour. A tool that only writes out an assessment made by humans does not pursue that purpose. Obligations for stand-alone Annex III high-risk systems apply from 2 December 2027 after the shift introduced by the digital omnibus package.

Does the works council have to be involved?

As a rule yes. The employer must inform the works council about the planning of work procedures including the use of artificial intelligence and consult it in good time. Where AI is used in selection guidelines, co-determination rights expressly apply, and for assessing AI the involvement of an external expert is deemed necessary.

Can a reference letter now be issued electronically?

Yes, but only with the employee's consent. Since the Fourth Bureaucracy Relief Act came into force on 1 January 2025, the reference may be issued in electronic form; before that it was excluded. Electronic form requires a qualified electronic signature, and without consent the written original remains the rule.