A New Study Found AI Is Quintupling Complaint Volume at Government Agencies. Your Contact Form, RFP Inbox, and Warranty Queue Are the Small-Business Version of the Same Problem.
Researchers documented 84 cases across 11 countries where AI-generated submissions overwhelmed public-service intake systems — UK housing complaints up nearly 3x, US CFPB complaints up 5x. Most of it isn't spam; it's real people using AI to clear the effort bar. Here's what that means for the intake forms your small business runs on.
Note: TechCrunch reported on this research September 10, 2026 — two days old as of writing. The underlying paper, “Characterizing Agentic Flooding of Government Services,” is a preprint set to be presented at the AI Ethics and Society conference; the dataset behind it (84 documented cases across 11 jurisdictions, spanning 2018–2025 volume data) is publicly posted.
Every one of this site’s pieces on AI agents so far has been about running them — sandboxing them, patching around them, protecting a network from one that misbehaves. This one’s different. It’s about what happens to the other side of your business — the inbox, the contact form, the intake queue — once enough of your customers, applicants, and vendors are running agents of their own to talk to you. A new study says that’s already happening to government agencies, at scale, and there’s no structural reason it stops at government.
What the study actually found
Researchers from the Centre for Digital Governance at the Hertie School in Berlin, the UK’s Cooperative AI Foundation, and the think tank GovAI examined 84 documented cases across 11 countries — Australia, Brazil, Denmark, Estonia, France, Germany, Japan, the Netherlands, Singapore, South Korea, the UK, and the US — where a public-service intake system saw a sharp, attributable jump in submission volume or complexity tied to AI use. They call the phenomenon “agentic flooding.”
The numbers, where they published specifics:
| Service | Change |
|---|---|
| UK Housing Ombudsman complaints | 2,600 (2022) → 7,000+ (2025) — nearly 3x |
| US Consumer Financial Protection Bureau complaints | ~5x growth since ChatGPT’s introduction |
| Brazilian judicial petitions | Documented surge, not separately quantified |
| German parliamentary petitions | Documented surge, not separately quantified |
Across the 84 cases in the dataset, 60% showed a clear rise in submission volume, 90% showed a rise in submission complexity (longer, more detailed, more legally structured filings), and 42% showed both at once. The researchers attribute 87% of the pattern to one simple, unglamorous cause: large language models can generate competent-looking text — an appeal, a complaint, a benefits application — cheaply and near-instantly, at a level of polish that used to require paying a lawyer, a consultant, or spending an evening you didn’t have. To collect the dataset itself, the researchers built an LLM-aided pipeline that scanned government sources across 13 domains in 12 countries, applying three inclusion standards before counting a case: a plausible AI cost-reduction mechanism, documented evidence of a demand-pattern change from an official source, and external attribution to AI use specifically — with human verification stages built in to catch false positives.
The part that should actually change how you think about this
Here’s the finding that makes this worth a dedicated piece instead of a footnote: the researchers explicitly found this mostly isn’t spam or bad-faith flooding. Lead researcher Chris Schmitz put it directly: “The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing.” The services seeing the biggest surges — tax appeals, court claims, welfare applications — are ones the paper describes as “financially attractive, but complex,” meaning there was always latent demand from people with a legitimate claim who never filed because the paperwork itself was the barrier. AI didn’t invent new claims. It removed the friction that used to suppress a chunk of real ones.
That reframe matters, because it means the honest business response isn’t “add a CAPTCHA and hope it goes away.” A CAPTCHA stops a bot; it does nothing to a real customer who used ChatGPT to draft a genuinely valid warranty claim, RFP response, or complaint, and is submitting it as themselves. The volume increase is coming from your actual audience getting more capable at making the same fundamentally legitimate asks, not primarily from an external adversary you can filter out.
Why this is a small-business story, not just a government one
Government intake systems got hit first for a structural reason: they’re often the most friction-heavy interfaces that exist — long forms, formal appeal language, legal thresholds — which made them the biggest single beneficiary of an AI tool that can generate that specific kind of text on demand. But nothing about the underlying mechanism is government-specific. Any small business with an intake process that used to be gated by effort is exposed to the identical dynamic, just with different names on the forms:
- RFP and quote-request inboxes. A contractor, agency, or service business that takes RFPs now faces the possibility of receiving a flood of AI-drafted, superficially thorough proposals or requests from parties who wouldn’t have bothered writing a real one by hand — some genuine, some fishing, and increasingly hard to tell apart at a glance.
- Warranty, return, and dispute claims. A well-written, detailed, policy-citing warranty claim used to be a costly-enough-to-produce signal that the claim was probably genuine. That signal is weakening — a customer can now generate an equally polished claim in thirty seconds regardless of how legitimate the underlying issue is.
- Job applications and cover letters. Already a widely reported problem: AI-generated cover letters and resumes tailored per posting, submitted at volumes no individual applicant could hand-write, straining small-business hiring pipelines that assumed a cover letter reflected real per-application effort.
- Support tickets and complaint forms. This site’s piece on why contact forms silently lose leads is about legitimate leads getting lost in the pipeline. This is close to the inverse problem: the pipeline filling up with a higher volume of longer, more polished, harder-to-triage submissions, some real, some not, at a rate the same small team has to somehow keep processing.
What made the paper’s numbers move, and why it’s not slowing down
The researchers note the growth pattern was largely flat before 2022, then accelerated as generative AI tools became broadly accessible, and — critically — show no sign of slowing as of the study’s most recent data. That’s consistent with what anyone watching AI tool adoption curves would expect: the capability that drives agentic flooding (cheap, competent text generation) has gotten more accessible, not less, since 2022, and the newest wave of AI agents can go a step further than drafting text — they can research the applicable rules, fill out the form fields themselves, and submit on a schedule, without a human doing each step.
That’s the genuinely new-frontier piece worth flagging as an experimental, forward-looking angle rather than settled fact: today’s documented flooding cases are mostly “a human used an AI tool to draft something, then submitted it themselves.” The version worth watching for is an agent handling the entire loop — drafting, submitting, and following up on a claim or application autonomously, at a volume and persistence no single human customer would sustain. Nothing in the current dataset proves that’s happening yet at scale outside government pilots, but the mechanism (agents with browser/form-filling capability, covered in this site’s piece on OpenAI’s computer-use launch) already exists. A small business’s RFP inbox or warranty-claim form is not a harder target than a government benefits portal — arguably it’s an easier one, since most small-business intake forms have no rate limiting, no identity verification, and no one watching submission-volume trends at all.
What the researchers recommend — and what actually translates to a small shop
The paper’s three recommendations were written for governments, but the underlying logic holds up at small-business scale with some translation:
- Audit your own exposure. The paper’s framework scores services on effort-to-submit versus expected benefit of submission — the combination that predicts flooding risk. Translate that to your business: which of your intake channels combine “low effort to submit now that AI drafts it” with “meaningful payout if approved” — a warranty claim, a dispute, a refund request, an RFP win? Those are your highest-risk channels, not your general contact form, which usually has low payout on either side.
Laid out as a grid, most small businesses have three or four intake channels that fall into genuinely different risk tiers, and it’s worth actually sorting yours instead of treating “forms” as one undifferentiated bucket:
| Channel | Effort to submit (post-AI) | Typical payout to submitter | Flooding risk |
|---|---|---|---|
| General contact / newsletter signup | Already near-zero | None or trivial | Low — nothing to game |
| RFP / quote request | Low (AI drafts convincingly) | High if it wins the job | High — worth a specificity filter |
| Warranty / return claim | Low (AI drafts convincingly) | Refund, replacement, or credit | High — worth a verification step |
| Job application | Low (AI drafts per-posting) | A job | High, but usually already triaged by hand |
| Client support ticket (existing customer) | Low | Usually just a fix, not a payout | Low — identity is already known |
The pattern in that grid is the paper’s finding in miniature: risk tracks payout, not raw submission volume. A channel can get busier without getting riskier if there’s nothing real to gain by gaming it, and a channel with real money or a real decision attached deserves scrutiny even if its volume hasn’t visibly spiked yet. 2. Identity, not friction, is the sustainable filter. The paper’s second recommendation — digital identity to enable rate limits without blocking legitimate access — translates directly: a logged-in customer account tied to a real purchase history, order number, or prior relationship lets you rate-limit and prioritize by known-real signals, instead of adding generic friction (CAPTCHAs, longer forms) that inconveniences your genuine customers roughly as much as it slows down anyone gaming the system. 3. Decide your policy before volume forces a bad one. The paper found only 56% of documented government cases had any explicit response at all, and most of those were limited pilots rather than real policy — meaning most institutions are reacting after the fact, under pressure, rather than deciding in advance how they’ll triage a flood. A small business is better positioned than a government agency to actually do this in advance: decide now whether AI-assisted submissions get treated identically to any other submission (the honest, Schmitz-aligned answer, since most are legitimate), or whether certain high-payout channels need a verification step added before volume forces a rushed decision.
A worked example: the RFP inbox
Say you run a five-person contracting or agency business that takes RFP submissions or quote requests through a web form. Six months ago, you got maybe 15 genuine inquiries a month, each one representing real time someone spent explaining their project. Today you’re getting 40, several of them template-flavored, generically detailed in a way that reads competent but oddly non-specific to your business — a pattern consistent with AI-assisted drafting rather than a sudden surge in real demand.
The over-reaction is adding a long, effortful qualifying form to filter people out — which, per the paper’s own findings, does little against AI-assisted submissions specifically (the AI happily fills out a long form too) while meaningfully discouraging your real prospects from finishing it. The better move: add one or two questions that are cheap for a real prospect who actually knows their own project to answer specifically and expensively generic for a templated submission to fake convincingly — a specific budget range tied to a specific timeline, or a direct question about a constraint unique to their situation. Then triage by response specificity, not by length or polish, since polish is now the cheapest signal there is.
Worth naming the actual cost here, since “we just get more emails now” undersells it: triage time is the bottleneck, not inbox space. If reviewing each RFP used to take ten minutes and volume nearly tripled, that’s not a minor inconvenience — it’s roughly two extra hours a week of someone’s time spent reading proposals that increasingly all sound similarly polished, which is exactly the condition under which a genuinely strong proposal starts blending into the noise instead of standing out.
There’s a second-order cost worth naming too: the same dynamic that makes a fake RFP response harder to spot on sight also makes a genuinely excellent real proposal harder to distinguish from a merely competent AI-assisted one, at a glance. Reviewers under time pressure with triple the volume tend to default to skimming for red flags rather than reading closely for genuine strength — which means the honest cost of agentic flooding isn’t just more spam to filter, it’s a real risk of your team’s attention quality dropping on the submissions that matter most, simply because there are three times as many of them competing for the same ten-minute review slot.
A second worked example: the warranty and returns queue
The RFP inbox is a service business’s version of this problem. A product-based small business — say, an eight-person outdoor-gear retailer selling through its own site and a couple of marketplaces — feels it on the other end, in warranty and return claims instead of quote requests.
Before AI-drafted claims became common, a warranty claim that cited the specific defect, referenced the actual product model and purchase date, and explained the failure mode in coherent, policy-aware language was a reasonably strong signal the claimant had a real issue — writing a convincing claim from scratch took enough effort that most people casually gaming the system didn’t bother, and the ones who did were usually easy to spot because the writing gave them away. That signal is close to gone now. A customer with a genuinely marginal case — a product that failed after visible misuse, or just outside the warranty window — can generate a claim indistinguishable in tone and structure from one written by a customer with an airtight case, in under a minute, at no cost.
The instinct is to tighten the warranty policy itself — shorter windows, more exclusions, more required documentation. That mostly punishes legitimate customers, exactly the failure mode the study’s own authors warn about, and does very little to the specific problem, since a longer documentation requirement is just as easy for an AI-drafted submission to fill out as a real one. What actually holds up is shifting the verification burden from language to evidence: require a photo of the defect and the product’s serial number or order ID as a submission gate, not an optional add-on. That’s cheap for a genuine claimant, who has the product in hand, and meaningfully harder to fabricate convincingly than another paragraph of well-structured prose — which is exactly the same principle as the RFP example’s “specific budget range” question, applied to a different channel. The lesson generalizes: wherever AI has cheapened the writing, the fix is to raise the cost of the evidence, not the cost of the words.
Quick answers
Should I just start rejecting anything that looks AI-written? No — per the study’s own core finding, that penalizes your legitimate customers roughly as much as anyone gaming the system, since most people using AI to draft a submission have a genuine underlying claim and are simply using a tool to write it up competently. The goal is triaging by verifiable specifics, not by writing style.
Are AI-detection tools a real fix here? Only partially, and with real false-positive risk. AI-detection accuracy on polished business writing is inconsistent, and a detector flagging a genuine customer’s AI-assisted warranty claim as “suspicious” creates exactly the kind of bad customer experience this piece is trying to help you avoid. Specificity-based triage (see above) is a more reliable signal than detection-tool output.
Does this mean I should require an account or login for every form? Not every form — only the ones tied to real payout (warranty claims, RFPs, disputes), per the paper’s own effort-versus-benefit framework. Adding login friction to a low-stakes newsletter signup or general contact form solves a problem that channel doesn’t have, and just loses you leads, which is the opposite failure mode covered in the contact-form piece.
Is there a way to actually measure whether this is happening to my business yet? Track submission volume and average length/complexity over time for your highest-payout intake channels, the same way the researchers did for government services. A sudden, sustained jump with no matching increase in your actual customer base or marketing spend is the same signal the paper used to flag a case for inclusion.
When this doesn’t apply to you
If your intake volume is low and personally handled anyway — a five-inquiry-a-week solo operation where you read and respond to every submission by hand — flooding at the scale the paper documents isn’t a near-term operational risk, though a single well-crafted, high-effort-looking submission is still worth verifying before you treat it as automatically credible.
If none of your intake channels have a meaningful payout attached — a basic newsletter signup form, a general “contact us” with no financial or approval stakes — you’re not the profile this paper is describing. Flooding concentrates where the paper says it will: effort-reducing tools meeting genuinely valuable, previously friction-gated outcomes.
If you already require account-based, verified submission for anything with real stakes (a client portal, an order-linked warranty system) — you’ve effectively already implemented the paper’s own recommended fix, ahead of the problem showing up.
Sources
All facts accessed September 12, 2026.
- Original reporting on the study’s release and researcher quotes — TechCrunch, “AI agents are flooding public services with new requests”
- Full methodology, dataset scope, and the three policy recommendations — arXiv, “Characterizing Agentic Flooding of Government Services”
- Additional reporting on complaint-volume statistics — Dataconomy, “Study Finds AI Linked To Surges In Government Complaints”
Bottom line
The researchers behind this study weren’t describing a spam problem; they were describing a friction problem that AI just solved for anyone with a legitimate claim and no time to write it up properly by hand. Government intake systems felt it first because they had the most friction to remove. A small business’s RFP inbox, warranty queue, or application pipeline runs on the exact same friction-gated logic, with no structural reason it’s immune. The move isn’t to build a wall against AI-drafted submissions — most of what’s coming through it is real. It’s to figure out, before volume forces the decision under pressure, which of your intake channels actually need a better filter than “longer form, please.”