Smartphone showing a sports betting app beside rising analytics charts illustrating AI behavioral targeting of gamblers

DraftKings built a machine learning model that does exactly what it was asked to do. According to the New York Times, the company trained it on customer betting records to identify gamblers most likely to lose, then targeted them with promotions designed to bring them back. The model works. Revenue goes up. And the people most likely to be re-engaged are the ones already gambling despite damage to their finances and relationships. No third-party data was needed. DraftKings did it with first-party data it already had.

This is the failure mode that should worry you, and not because you run a betting app. An AI system can hit its KPI perfectly and still do something your organisation would never approve if a human proposed it out loud. Below, what that looks like in an operations setting, and how to design it out before deployment.

A Model That Works Exactly as Designed, And That’s the Problem

There is no bug here. No data leakage, no drift, no bad labels. The engineering is sound, the predictions are accurate, and the business case closes. The choice of what to optimize for was the failure, and that choice was made by people, not by the model.

That distinction matters because it moves the risk out of your data science team and into your management meeting. As EFF’s Devanshi Nishar points out, AI also makes this harder to audit after the fact: “Because AI operates as a black box, the humans building the models can rarely predict which data points are the most useful to the AI.”

So if you approve a KPI before mapping its second-order consequences, you have already approved the outcome. The model just gets there faster.

Dashboard screen showing AI behavioral targeting scores ranking gamblers by predicted spending

What DraftKings Actually Built, and Why First-Party Data Changes the Argument

Strip away the outrage and the architecture is unremarkable. Betting records go in as training data. A classifier ranks users by probability of losing. A re-engagement campaign fires at whoever scores highest. Any operations team that has built a churn model or a maintenance prioritisation model has built the same shape of thing.

That is the uncomfortable part. Nothing in the pipeline looks different from work you would sign off on this quarter.

The training data was the customer relationship itself

DraftKings did not scrape anything or buy a data broker feed. It used records generated by people using the product as intended, which is the cleanest, most consented data a company can have. Every bet placed was a labelled example.

Apply that to your own operation. Your maintenance logs, operator performance records, supplier scorecards and warranty claims are all high-quality training data produced by the ordinary business relationship. That is exactly why they carry weight. A model trained on how your people and partners actually behave can be pointed at outcomes that serve them or outcomes that quietly work against them, and the data looks identical either way.

Why ‘we only use our own data’ is not a governance control

Most AI governance checklists in circulation ask where the data came from. Third-party purchase, consent, retention, deletion. DraftKings passes all of it. As EFF notes, the company appears to be using solely first-party data, and the organisation argues that policy fixes aimed only at third-party data sharing would not be enough to prevent this.

So data provenance answers a legal question, not an ethical one. The thing that determined the harm here was the target variable, not the source. If your review board approves models on the basis of data lineage and says nothing about what the model is optimising toward, you have built a compliance artefact rather than a control.

Write the objective function into the approval document. Make someone senior sign their name next to it.

How AI Amplifies a Twenty-Year-Old Advertising Problem

Targeting people based on collected data has been standard practice since the early 2000s. What changed is throughput and appetite. A model can score millions of records in the time a marketing team used to spend arguing over segment definitions, and it will do it again tomorrow with fresh data.

EFF’s argument is that AI magnifies harms that already existed rather than creating new ones. Behavioral advertising already rewarded companies for hoarding data. Training and refining models raises the reward, because more training data usually means a sharper classifier. The incentive runs one direction only.

The black box creates a permanent appetite for more data

Here is the mechanism that operations leaders underrate. When you cannot explain which features are carrying the prediction, you cannot justify dropping any of them. So you keep every field, add more, and retain it longer, on the theory that the model might find something. Collection becomes the default rather than a decision someone signs off on.

That is a governance problem dressed up as an engineering one. A team that cannot state which inputs matter also cannot answer a regulator, an auditor, or a customer asking why a given data point was collected. Feature importance analysis and a documented minimum viable dataset are not academic exercises. They are what lets you say no to the next collection request.

The downstream consequence is where this gets genuinely serious. EFF notes that data gathered for ad placement gets sold onward to insurance companies, banks, and government agencies including CBP. Earlier this year ICE published a Request for Information “seeking information to better understand how the industry’s commercial Big Data and Ad Tech providers can directly support investigations activities.”

You may never sell a record. But data you collected because a model might use it still sits in your systems, subject to subpoena, breach, and acquisition. Every field you retain without a stated purpose is a liability you are holding on someone else’s behalf.

Diagram of three data streams feeding an AI behavioral targeting engine that profiles website visitors

The Same Failure Mode in a Manufacturing or Quality Context

Where optimization targets diverge from stated intent

A scheduling model told to maximise line throughput will find inspection depth. It will not announce that it is trading quality assurance for units per hour. It will just route more batches through the fastest path, and the fastest path is usually the one with the lightest checks. Your OEE dashboard will look excellent right up until the field returns arrive.

Supplier scorecards fail the same way. Train one on your ERP records and it learns to punish vendors with messy data reporting, not vendors with actual defects. The small supplier who emails a spreadsheet gets downgraded. The large one with clean EDI integration and a worse defect rate gets promoted.

Workforce analytics is the third common case. A model trained on historical shift outcomes absorbs whatever your past scheduling decisions encoded, including the ones nobody would defend out loud, then projects them forward as an objective recommendation.

Questions to ask before a model reaches production

DraftKings only used data it collected directly from its own users, no third-party purchase required. That detail should land hard, because it means your existing MES, ERP, and QMS data is already enough to build something you would not want to explain to a customer. Governance has to happen at the objective-setting stage, not at the data-acquisition stage.

Run these before anything gets deployed:

  • What gets sacrificed to improve this metric? Name it explicitly. If nobody can answer, the model has not been reviewed.
  • Who is at the bottom of the ranking, and why? Pull the ten lowest-scored suppliers, operators, or batches and check the reason. Data quality artefacts show up fast.
  • What guardrail metric stays fixed? Throughput targets need a floor on inspection coverage that the model cannot move.
  • Would you defend this logic in an audit? If the optimisation target sounds bad said plainly, it is bad.

Write the answers down. Unrecorded intent is how these programs drift.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

Building AI Programs That Survive Public Scrutiny in 2026

EFF makes a point worth sitting with: policy fixes aimed at third-party data sharing and selling would not have stopped this model. The training data never left the company. Any rule written around brokers and resale misses a system built entirely on the customer relationship.

That means the constraint has to come from inside. If you are waiting for a regulator to tell you which optimization targets are acceptable in your plant, you will be waiting past the point where it matters.

The governance checkpoint most AI projects skip

Most AI model governance in manufacturing covers data quality, model performance, and access control. Almost none of it covers the objective function. Nobody signs off on what the model is being paid to maximise, which is the only decision that determines whether the output is defensible.

Three things make a program hold up. Write the optimization target in plain language and have it reviewed by someone who did not build the model. Log the harm scenarios next to the expected ROI in the same document, so both travel together. Define the kill criterion before launch, including who has the authority to pull the model and what number triggers it.

Why reputational risk belongs in the business case

A model that gets switched off after a bad news cycle returns nothing. You paid for the data work, the integration, the change management, and the training, and then you wrote it all off. Worse, the next three AI proposals in your organisation get slower approval because someone remembers.

Governed models compound instead. They stay in production, accumulate operating history, and build the internal credibility that makes the next deployment easier to fund. That is the actual return on responsible AI deployment, and it shows up as deployment velocity rather than a line in a single business case.

AI transformation and AI restraint are the same discipline. Knowing which models not to build is what earns you the right to keep building.

Source: eff.org

Leave a Reply