Twee jaar lang ging het gesprek over AI over de vraag hoe snel machines werk kunnen maken: code, teksten, analyses, besluiten. Deze week schoof in drie berichten die elkaar niet kennen hetzelfde probleem naar voren: het maken gaat inmiddels zo snel dat het beoordelen de bottleneck wordt. Een peiling onder developers geeft het verschijnsel een naam, een boete van 825 miljoen euro geeft het een prijs en de grootste Deense pre-seedronde ooit geeft het een verdienmodel.For two years the conversation about AI was about how fast machines can produce work: code, texts, analyses, decisions. This week three reports that know nothing of each other pushed the same problem forward: making things has become so fast that reviewing them is becoming the bottleneck. A survey of developers gives the phenomenon a name, a fine of 825 million euros gives it a price and the largest Danish pre-seed round ever gives it a business model.
Developers lopen vast op hun eigen tempoDevelopers are getting stuck on their own pace
Een kleinschalige peiling van leerplatform Coddy onder 305 developers, opgetekend door ZDNet en AG Connect, vat de stemming samen: 80 procent zegt dat het werken met AI meer aanvoelt als een afhankelijkheid, een verslaving, dan als een voordeel. Tegelijk zegt 74 procent dat intensief gebruik de kans op loonsverhoging of promotie vergroot en ziet 51 procent de kans op een burn-out toenemen. De peiling is klein en komt van een belanghebbende partij; als los cijfer verdient hij terughoudendheid.A small-scale survey by learning platform Coddy of 305 developers, recorded by ZDNet and AG Connect, sums up the mood: 80 per cent say that working with AI feels more like a dependency, an addiction, than an advantage. At the same time 74 per cent say that intensive use increases the chance of a pay rise or promotion and 51 per cent see the risk of burnout growing. The survey is small and comes from an interested party; as a standalone figure it deserves caution.
Interessanter dan de percentages is het begrip dat erbij hoort: „verification debt”, een achterstand in het aantonen dat werk klopt. Code is snel gegenereerd, maar de controle of het resultaat correct, veilig en beheersbaar is kost meer tijd, en die achterstand stapelt zich op. Het verschijnsel is breder zichtbaar dan één peiling. Vorige week nog stelde Microsoft een grote Exchange-update uit omdat AI-systemen meer kwetsbaarheden vinden dan menselijke beoordelaars kunnen verwerken, zo meldden wij in het weeknieuws van week 35. De generatiekant schaalt; de beoordelingskant niet vanzelf.More interesting than the percentages is the term that comes with them: ‘verification debt’, a backlog in showing that work is correct. Code is generated quickly, but the verification of whether the result is correct, secure and manageable takes more time, and that backlog piles up. The phenomenon is visible more widely than one survey. Only last week Microsoft postponed a major Exchange update because AI systems find more vulnerabilities than human reviewers can process, as we reported in the week 35 news round-up. The generation side scales; the review side does not do so by itself.
Wie de beoordeling overslaat, kent nu de prijsWhoever skips review now knows the price
Wat er gebeurt als een organisatie het beoordelen wegautomatiseert, kreeg onlangs een bedrag.What happens when an organisation automates review away was recently given a figure.
De Autoriteit Persoonsgegevens legde Uber die boete op omdat het bedrijf tussen 2018 en 2022 chauffeursaccounts automatisch blokkeerde op basis van rijgedrag, klantbeoordelingen en vermeende fraude, zonder dat een mens de besluiten beoordeelde; ook wisten chauffeurs onvoldoende wat er gebeurde. De grondslag is artikel 22 van de AVG, dat volledig geautomatiseerde besluiten verbiedt die iemand juridisch of anderszins aanmerkelijk raken.The Dutch Data Protection Authority imposed that fine on Uber because between 2018 and 2022 the company automatically blocked driver accounts based on driving behaviour, customer ratings and suspected fraud, without a human reviewing the decisions; drivers also knew too little about what was happening. The basis is Article 22 of the GDPR, which prohibits fully automated decisions that affect someone legally or otherwise significantly.
Alleen het bedrag is uitzonderlijk, de regel niet. TheAIDaily wijst erop dat dezelfde eis geldt voor het fraudefilter dat een account blokkeert, het sollicitatiesysteem dat afwijst op een score en de webshop die achteraf betalen weigert op een kredietscore. En de lat ligt hoger dan een formaliteit: iemand die op akkoord klikt omdat het systeem dat adviseert, is geen waarborg. Menselijke beoordeling moet het besluit kunnen tegenhouden, anders telt ze niet.Only the amount is exceptional, the rule is not. TheAIDaily points out that the same requirement applies to the fraud filter that blocks an account, the application system that rejects on a score and the web shop that refuses deferred payment on a credit score. And the bar is higher than a formality: someone clicking approve because the system advises it is no safeguard. Human review must be able to stop the decision, otherwise it does not count.
De markt bouwt bedrijven om de beoordeling heenThe market is building companies around review
Waar een knelpunt zit, ontstaat een markt. De oprichters van de Deense neobank Lunar haalden deze week 8,2 miljoen euro pre-seed op, volgens Sifted-data de grootste ronde van Denemarken tot nu toe, voor Repodo: een kantoor voor wettelijke controles dat vanaf de start rond AI is opgezet. Het platform automatiseert de uitvoering, waaronder dataverzameling, aansluitingen, documentatie en transactieanalyse, terwijl bevoegde accountants verantwoordelijk blijven voor oordeelsvorming, toezicht en ondertekening. Repodo begint in Denemarken bij het mkb en wil daarna land voor land door Europa uitbreiden.Where there is a bottleneck, a market emerges. The founders of Danish neobank Lunar raised 8.2 million euros pre-seed this week, according to Sifted data the largest round in Denmark to date, for Repodo: a statutory audit firm set up around AI from the start. The platform automates the execution, including data collection, reconciliations, documentation and transaction analysis, while licensed accountants remain responsible for forming the opinion, oversight and signing off. Repodo starts in Denmark with SMEs and then wants to expand through Europe country by country.
De taakverdeling is het bericht. Investeerders steken hun geld niet in een systeem dat het oordeel overneemt, maar in een inrichting waarin de machine het verzamelwerk doet en de mens aantoonbaar eindverantwoordelijk blijft. Dat is dezelfde les die Nvidia's harnas-onderzoek vorige week voor agents trok: niet het model maakt het verschil, maar de sturing, begrenzing en controle eromheen.The division of labour is the news. Investors are not putting their money into a system that takes over the judgement, but into a setup in which the machine does the gathering work and the human demonstrably remains ultimately responsible. That is the same lesson Nvidia's harness research drew for agents last week: it is not the model that makes the difference, but the steering, boundaries and verification around it.
Het gevolg: organiseer de beoordeling als volwaardig werkThe consequence: organise review as work in its own right
De drie berichten wijzen dezelfde kant op. Wie AI het maakwerk geeft, moet drie dingen regelen: capaciteit, verantwoordelijkheid en aantoonbaarheid.The three reports point in the same direction. Whoever gives AI the production work must arrange three things: capacity, responsibility and demonstrability.
Terug naar het begin: het maken is niet meer het probleem. De organisaties die de komende jaren het verschil maken, zijn niet de organisaties die het meest genereren, maar de organisaties die kunnen aantonen dat wat ze genereren klopt.Back to the beginning: making things is no longer the problem. The organisations that make the difference in the coming years are not the organisations that generate the most, but the organisations that can show that what they generate is correct.
Lees ook.Also read. Waarom het harnas, niet het model, bepaalt of een agent presteert: Het harnas, niet het model…Why the harness, not the model, determines whether an agent performs: The harness, not the model…
Hoe wij digitale collega's omgeven met een harnas van vakkennis, controle en security by design: de softwarepagina.How we surround digital colleagues with a harness of domain expertise, control and security by design: the software page.
BronnenSources
- AG Connect - 80% van de developers verslaafd aan werken met AI
- TheAIDaily - 825 miljoen boete voor Uber: zo check je of jouw AI-besluit een mens nodig heeft
- Sifted - Lunar founders raise €8.2m to launch AI-native audit startup Repodo · Finextra · Tech.eu
- AI-weeknieuws week 35 - Exchange-uitstel en het harnas-onderzoek