Dutch police ran a predictive policing algorithm for a decade with zero proof it worked. Your AI vendor might be doing the same.
elcome to issue seventeen. This week's stories share a theme: AI systems that were never actually tested against the outcomes they claimed to improve. The Dutch police quietly scrapped their Crime Anticipation System in early 2026 after an internal report concluded it had no measurable effect on crime, despite a decade of deployment. Meanwhile, AlgorithmWatch found four commercial chatbots steering pregnant users toward anti-abortion advocacy sites without disclosing the source, which is exactly the kind of opacity the EU AI Act's transparency rules exist to catch.
Let’s go.
yours, Flux

Flux Weekly is a 6-minute briefing for people who have to actually make AI work in Europe. Sole traders to enterprise, one issue every Friday morning.

- New We added a vendor evidence checklist to the Flux compliance toolkit this week, prompted by the Dutch predictive policing story: twelve questions to ask before renewing an AI contract.
- Updated The high-risk AI category explainer in our resource library now includes live facial recognition as a worked example, reflecting the Nottinghamshire Police situation.
- ICYMI Last week's piece on Flock Safety is worth re-reading alongside this week's Dutch policing story: both are case studies in what happens when AI procurement skips outcome validation entirely.
Dutch Police Scrapped Their Predictive Policing Algorithm After a Decade, Having Never Proven It Worked

A decade. No evidence. The Netherlands' Crime Anticipation System was shut down in early 2026 after an internal report delivered a damning verdict: the system's predictive capabilities were, in the report's own words, extremely limited, and the police were never able to demonstrate any measurable impact on reducing crime. Ten years of deployment, zero validated outcomes. If that sentence does not make you want to audit your own AI vendors, read it again.
This is the high-risk AI problem in miniature. Predictive policing sits squarely in the category of high-risk AI under the EU AI Act, which requires fundamental rights impact assessments and ongoing monitoring of real-world performance. The Dutch case is a textbook example of what happens when neither of those things occurs. The system was not necessarily malicious. It was simply never tested, never validated, and never discontinued until someone finally wrote it all down.
Does your AI inform a decision that affects a person's job, credit, education, or essential service?

- ✓Dutch police Crime Anticipation System discontinued after internal report found no measurable crime-reduction impact across a decade of use.
- ✓AlgorithmWatch experiment found four commercial chatbots in Europe directing users to anti-abortion advocacy content without clear source disclosure, raising transparency concerns under AI Act Article 13.
- ✓Civil society coalition including EFF, Liberty, and Big Brother Watch pressed Nottinghamshire Police to halt live facial recognition rollout, a technology classed as high-risk under the AI Act.
- ~US Ninth Circuit ruled that platforms hosting user speech must endure lengthy, costly lawsuits before Section 230 dismissal, raising content-moderation costs globally.


- 1AlgorithmWatch CAS InvestigationCase study
The full AlgorithmWatch report on the Dutch Crime Anticipation System, covering its decade-long run and the internal review that ended it.
Why we like it. It is the clearest real-world template for what a failed AI monitoring regime looks like, and what your own audit should be designed to catch.
- 2EU AI Act Article 9 Checklist (CEPS)Compliance
The Centre for European Policy Studies published a plain-language breakdown of Article 9 risk management obligations for high-risk AI systems.
Why we like it. The Dutch CAS had none of the ongoing monitoring Article 9 requires. This checklist tells you what the law expects you to have in place.
- 3AlgorithmWatch Chatbot Health Bias Report

Ten years is a long time to trust something you never tested
By John Ferguson
The Dutch predictive policing story hit differently this week. Not because it is shocking, but because it is familiar. A tool gets adopted, budgets get committed, workflows get built around it, and the question of whether it actually works quietly stops being asked.
That is not unique to policing. I have spoken to operators running AI tools in customer service, credit assessment, and HR who have never once compared outcomes before and after deployment. The tool exists. It produces outputs. Everyone assumes that is the same as it working.
The EU AI Act is going to force that question back onto the table, at least for high-risk systems. Article 9 and the post-market monitoring requirements are not just bureaucratic boxes. They are a formal insistence that evidence of performance is not optional.
If I am honest, the December 2027 deadline feels abstract until you read a story about a decade wasted on a system no one ever validated. Then it feels urgent. Start the conversation with your vendor now, while you still have time to do something about the answer.
John Ferguson · Founder, Agentic Fluxus

Short answer.Recruitment scoring sits firmly in the high-risk category. You have time, but not much slack. The immediate priority is a documented risk assessment and a retrospective look at whether outcomes have been monitored at all since launch. The Dutch policing story is a cautionary tale: systems run for years without evidence review are exactly what auditors will look for first.
How does your organisation currently validate that an AI tool is actually delivering the outcomes the vendor promised?

The Dutch Crime Anticipation System ran from roughly 2016 to early 2026 with no validated evidence that it reduced crime. An internal report finally forced the shutdown, which raises the obvious question: who was reviewing the evidence in years two through nine?
AlgorithmWatch tested four commercial chatbots with questions about pregnancy options and found repeated referrals to pro-life advocacy sites, often without any disclosure that the source was not a health authority. Two of the four chatbots failed to distinguish official medical guidance from campaign content.

