Skip to content
Confir.
Risk Classification

AI That Evaluates Evidence Reliability: High-Risk Under Annex III Point 6(c)

High-Risk Use Case30 July 2026· 10 min read

AI evaluating evidence reliability in criminal cases is high-risk under Annex III point 6(c). Scope, actor test, obligations, 2 August 2026 deadline caveat.

A forensic tool scores how likely a seized video has been manipulated, and a police unit uses that score to decide whether the exhibit is worth sending to an expert. That tool is high-risk. It evaluates the reliability of evidence in a criminal investigation, by or on behalf of a law-enforcement authority, which is exactly the function captured by Annex III point 6(c) of Regulation (EU) 2024/1689. This page covers that one limb — evidence reliability — and how it differs from the other point 6 functions and from judicial AI.

The stakes are concrete. A breach of the high-risk obligations carries a maximum fine of €15 million or 3% of total worldwide annual turnover, whichever is higher, under Article 99(4). Stray into a prohibited practice and the ceiling rises to €35 million or 7% under Article 99(3).


What Annex III point 6(c) actually targets

The statutory wording, verbatim where possible

The limb covers AI intended to be used by or on behalf of law-enforcement authorities — or by Union institutions, bodies, offices or agencies in support of them — to evaluate the reliability of evidence in the course of investigation or prosecution of criminal offences. That is Annex III point 6(c) of Regulation (EU) 2024/1689. The sub-letter ordering inside point 6 should be verified against the consolidated text before you rely on it in a formal filing; the function described here is the operative test, not the letter.

What 'reliability of evidence' covers in practice

Evaluating reliability means scoring, ranking, authenticating, or flagging the trustworthiness of evidentiary material. In practice this captures forensic image, audio, and video authenticity detectors; deepfake and manipulation flaggers; fingerprint or ballistics confidence-scoring; and statement-consistency analysis applied to the evidence itself rather than to a person.

The output is what matters. The system conditions whether and how a piece of evidence is treated as credible in an investigation or prosecution. That is where the fundamental-rights stakes attach, and why the classification is strict.

Contrast adjacent tooling that is not this limb. Pure document management, transcription without analysis, and chain-of-custody logging that records but does not assess reliability all sit outside point 6(c). The trap is assuming that anything touching evidence is high-risk — it is the act of assessing trustworthiness, not mere handling, that triggers the classification.


The actor test: only law-enforcement use triggers point 6

Provider vs deployer in a forensic-AI supply chain

Point 6 is gated by who uses the system, not only by what it does. The system must be intended for use by, or on behalf of, law-enforcement authorities, or by Union institutions, bodies, offices or agencies supporting them. A vendor that builds and sells a forensic authenticity tool to a police force is the provider; the police force is the deployer.

"On behalf of" captures outsourced forensic laboratories and contractors acting for the authority. The classification follows the law-enforcement purpose, not the corporate identity of the operator — a private lab analysing exhibits for a prosecutor is squarely inside scope.

Under Article 25, a deployer, distributor, or importer that puts its name or trademark on the system, substantially modifies it, or repurposes it to this high-risk use becomes a provider and inherits the full provider obligation stack. Rebranding a tool does not relocate the obligations; it relocates the provider.

When the same tool falls outside point 6

The same algorithm sold to a newsroom for fact-checking is outside Annex III point 6 entirely, because the actor test fails. That does not make it unregulated: an authenticity tool that generates or labels synthetic media may attract Article 50 transparency duties on its own track. But the high-risk stack described below does not bite unless the law-enforcement actor test is met.


Why it is high-risk: fundamental-rights stakes in justice

The harm model: false-positive and false-negative authentication

Evidence reliability sits upstream of charging and conviction decisions. An error there propagates into the presumption of innocence, the right to a fair trial and an effective remedy, and non-discrimination — Articles 47, 48, and 21 of the EU Charter of Fundamental Rights. Two failure modes are rights-affecting and the risk-management system must address both: false authentication, treating manipulated evidence as genuine, and false rejection, discarding genuine evidence as suspect.

Why Article 6(3) rarely rescues an evidence-reliability system

Article 6(3) lets a provider treat an Annex III system as not high-risk where it poses no significant risk to health, safety, or fundamental rights — for instance, a narrow procedural task that does not influence assessments about specific persons. That filter is very hard to claim here. Any system whose output conditions the evidentiary basis of a criminal case influences an assessment about a specific person, by definition. A provider that nonetheless claims the exemption must document the analysis rigorously and still register under Article 49.

The Article 5 boundary for this limb differs from the offending-risk limb. Unlike point 6(d), there is no profiling-based prohibition specific to evidence reliability. But the system must still not stray into prohibited practices — for example, Article 5 biometric categorisation or emotion inference — when applied to people connected to the evidence.


How point 6(c) differs from the other point 6 limbs and from judicial AI

Point 6 limbs side by side

LimbWhat it assessesObjectArticle 5 boundaryCovered on Confir
6(a)Risk of becoming a victim of crimePersonNo limb-specific prohibitionLaw-enforcement overview
6(b)Polygraph / emotional-state detectionPersonWatch Art 5 emotion/biometric limitsLaw-enforcement overview
6(c)Reliability of evidenceThing (the evidence)No profiling-only banThis article
6(d)Offending / re-offending riskPersonArt 5(1)(d) bans profiling-only predictionCrime-prediction AI
6(e)Profiling of natural persons in detection/investigationPersonProfiling triggers high-risk; no carve-outLaw-enforcement overview

Point 6(c) versus point 8(a): the investigation/adjudication line

The core distinction inside point 6 is object. Crime-prediction and re-offending risk under point 6(d) assess a person's likely future conduct; evidence reliability under point 6(c) assesses a thing's trustworthiness. That difference drives both the classification logic and which prohibition risks apply — the profiling-only ban in Article 5(1)(d) attaches to 6(d), not to 6(c).

Distinguish too from Annex III point 8(a), judicial AI. Point 8(a) covers a court or alternative-dispute-resolution body's own analysis of facts and law and the application of law to facts — the adjudication phase. Point 6(c) is the investigation and prosecution phase, by or for law-enforcement authorities, before the matter is a judicial determination. The same exhibit may pass from one regime to the other; each carries its own classification and obligations. For the limbs this page does not cover in depth, follow the dedicated siblings linked below.


The obligation stack for providers and deployers

Provider obligations (Articles 9–15, 43, 47–49)

A stand-alone Annex III system takes the full high-risk stack directly. The provider must operate a risk-management system (Article 9), data governance (Article 10), technical documentation (Article 11 and Annex IV), transparency to deployers (Article 13), human oversight by design (Article 14), and accuracy, robustness, and cybersecurity (Article 15).

Conformity assessment runs through internal control — Article 43 with Annex VI is the standard route for stand-alone Annex III systems, no notified body required unless the system is embedded in an Annex I product. The provider then issues an EU Declaration of Conformity (Article 47), affixes the CE marking (Article 48), and registers in the EU database (Article 49). Point 6 systems register in the non-public section, accessible to competent authorities but not to the public.

Deployer obligations and the Article 27 FRIA

Deployer duties under Article 26 include using the system within the provider's instructions, ensuring human oversight, keeping operational logs for the minimum retention period, monitoring operation, and notifying serious incidents. The provider in turn maintains post-market monitoring (Article 72) and reports serious incidents (Article 73).

Article 27 adds a Fundamental Rights Impact Assessment for public-authority deployers — and private bodies providing public services — before deployment. For evidence-reliability AI, the FRIA centres on the right to a fair trial, the presumption of innocence, and non-discrimination, and documents the mitigation measures and residual risks.

Penalty tiers

High-risk breaches reach €15 million or 3% of total worldwide annual turnover, whichever is higher (Article 99(4)). Prohibited-practice breaches reach €35 million or 7% (Article 99(3)). Supplying incorrect or misleading information to authorities sits in the third tier at €7.5 million or 1% (Article 99(5)). Under Article 99(6), SMEs and start-ups are capped at the lower of the percentage or the fixed amount.


The deadline note — plan against 2 December 2027

Statute date and the adopted deferral

The statute originally set the high-risk obligations for Annex III systems (Article 6(2) systems) to apply from 2 August 2026. The Digital Omnibus — passed by the European Parliament on 16 June 2026 and adopted by the Council on 29 June 2026 — defers this to 2 December 2027, pending Official Journal publication as a formality.

Plan against 2 December 2027. A blanket "stop the clock" was rejected, and not everything is delayed.

What is already in force

The Article 5 prohibitions have applied since 2 February 2025, and the Article 4 AI-literacy duty since 2 February 2025; neither was deferred. Both stay relevant if an evidence-reliability deployment touches people in a way that brushes against a prohibited practice.


Worked example: a forensic-AI vendor and a regional police force

Step-by-step classification

EvidIntel GmbH, a roughly 70-person forensic-software vendor, sells an audio/video authenticity tool that scores the likelihood that a media file has been manipulated. A regional police forensic unit deploys it during investigations to flag exhibits for expert review.

Work the classification in order. The law-enforcement actor test is met — the deployer is a police forensic unit. The function is evaluating the reliability of evidence, which is Annex III point 6(c), so the system is high-risk under Article 6(2). The Article 6(3) filter is inapplicable, because the output conditions the evidentiary basis of a criminal case and therefore influences an assessment about a specific person.

Who does what, and by when

EvidIntel (provider). It compiles an Annex IV technical-documentation pack with performance metrics disaggregated by media type and a register of known failure modes, runs an Article 9 risk analysis of false-authentication and false-rejection harms, completes internal conformity assessment, issues a Declaration of Conformity, and registers in the non-public EU database.

The police unit (deployer). It conducts an Article 27 FRIA centred on fair trial and the presumption of innocence, builds a human-oversight protocol so the score never substitutes for a forensic examiner's conclusion, keeps operational logs, and registers the deployment in the non-public database.

Timeline. Both plan to be compliant by 2 December 2027, the adopted Digital Omnibus deadline, pending Official Journal publication as a formality. Full stop.


How Confir helps

Confir's classification engine runs the system's intended purpose and functional description through the Annex III point 6 logic — separating the 6(c) evidence-reliability limb from 6(d) offending-risk and from point 8(a) judicial use — and returns the same documented finding every time, with the rule that fired shown in plain language. The engine is deterministic and rule-based — no model inference, no hallucination — and reproducibility is the point an audit needs.

For public-authority deployers, Confir scopes the Article 27 FRIA around the Charter rights most at stake in evidentiary contexts and produces a registration-ready record for the non-public EU database. For providers, it structures the Annex IV technical-documentation pack and the Article 9 risk-management record so the false-authentication and false-rejection harm analysis is captured methodically.


Manage your EU AI Act compliance in one place

Confir automates risk classification, technical documentation, and audit trails for any company. No consultants. No 6-month projects. 14-day free trial.

Start free trial →

Keep reading