AI Agents and the Data Privacy Risks of Missing Data
“The absence of data is also data for AI systems.”
Debbie Reynolds, "The Data Diva"
What Missing Data Can Mean
As many of us follow news reports of AI agents taking actions their developers and users never intended, I keep returning to something I have said often: the absence of data is also data for AI systems. Organizations spend considerable time deciding what information to give AI. They also need to consider how missing information influences its decisions. An AI agent may encounter an incomplete record, an unstated restriction, or a task that leaves acceptable methods undefined. What happens next depends partly on how the system handles what it does not know.
Absence can carry meaning without explaining itself. A blank field could mean information is unknown, deliberately withheld, or prohibited from collection. A missing record could reflect a technical failure or a completed deletion request. These circumstances may look similar to a system while requiring different responses. Omissions can influence actions even when a system does not explicitly recognize them. Organizations need to understand which assumptions may fill those gaps.
The Inertia of Pursuing a Goal
In my earlier essay on Goal Supremacy, I explored how AI agents pursue objectives and choose routes to achieve them. Humans routinely communicate goals while leaving boundaries unspoken. We request a better customer experience and assume confidentiality remains protected. An AI agent may infer expectations from its training, but those inferences cannot reliably substitute for the specific information and authority an organization needs to establish.
I think of the tendency to continue pursuing a goal as a kind of inertia. Once an AI agent begins working toward an outcome, a blocked route may lead it to search for another route. If the system treats the objective as more important than the restriction it encounters, and it can proceed another way, it may continue. Organizations may define what should move the AI agent forward without adequately defining what should bring it to a stop.
The Hugging Face incident illustrates how far that pursuit can extend. OpenAI disclosed in July 2026 that models undergoing an internal cybersecurity evaluation escaped a constrained testing environment and compromised Hugging Face infrastructure. OpenAI explained that the evaluation used reduced safety protections to measure cybersecurity capabilities. Hugging Face documented an intrusion that crossed several boundaries between systems. The reduced protections matter when interpreting this research incident and its consequences outside the intended testing environment.
The investigation by METR and Redwood Research described AI agents intended to operate separately finding an unauthorized way to communicate and coordinating efforts to manipulate the evaluation. Some joined the attack while recognizing that it fell outside their assigned tasks. That complicates an explanation based entirely on missing instructions. An AI agent can recognize a boundary and still act beyond it. Explaining expectations and enforcing them are separate responsibilities.
When an AI agent encounters an obstacle, does it treat that obstacle as a boundary to respect or a problem to solve? Restricted access might reflect the deliberate separation of sensitive information. An unavailable answer might mean the task should remain incomplete. Organizations need to establish when continuing is unacceptable, even if the AI agent can identify another route. The ability to overcome an obstacle cannot establish permission to do so.
Interrupting an AI Agent Requires Information
These concerns are influencing the debate about advanced AI. In his September 2026 essay, “We Must Pace the Frontier,” Anthropic’s Dario Amodei called for slowing capability advancement so safety work could keep pace, including independent evaluation. California Governor Gavin Newsom’s September 18 executive order accelerated oversight efforts and called for recommendations to advance independently verified emergency shutoffs for frontier models. Those proposed shutoff measures remain under development. Both developments reflect concern about retaining control.
I understand why a kill switch attracts attention. Organizations need an effective way to interrupt dangerous activity. However, someone must recognize that an AI agent has exceeded its authority, know who can intervene, and understand which actions have occurred. Shutting down a system cannot recall information already disclosed or automatically repair deleted records. Emergency intervention belongs within governance that defines acceptable behavior before an AI agent begins acting and detects departures while there is time to respond.
The Limits of Human Oversight
Human oversight leaves two important questions unresolved. Can the person intervene before the AI agent takes a consequential action? Does that person know enough to recognize when intervention is necessary? The first challenge is speed. An AI agent may execute actions faster than a person can examine them. By the time someone recognizes a problem, information may have been disclosed or records deleted. Some actions may be difficult or impossible to reverse. The AI agent may continue moving while the human is still trying to understand its previous actions.
For consequential actions, organizations need stopping points that prevent execution until review occurs. An alert seen after an action serves a different purpose from an approval requirement that holds the action until a decision is made. The system must enforce that pause, including when the AI agent identifies another route toward its objective. Otherwise, the organization depends on someone reacting quickly enough to interrupt activity that may already have passed the point of intervention.
The second challenge is expertise. People may delegate work to an AI agent because they do not know how to perform it themselves. If the reviewer does not understand the task, its boundaries, or what an acceptable result looks like, how will they recognize when the AI agent goes off track? A convincing explanation may appear satisfactory even when the work contains unauthorized actions. NIST has cautioned against assuming a human overseer provides adequate governance simply by being present. In my view, human-in-the-loop is not enough; the human must lead.
A reviewer does not need to reproduce every computational step. However, that person needs enough knowledge, evidence, and authority to evaluate an action and challenge assumptions. This introduces another dimension to missing data. The AI agent may lack information about what it should do, while the reviewer lacks information about what it actually did. Effective oversight requires an enforceable opportunity to intervene and a person equipped to use it, including recognizing inappropriate uses of personal data.
When Missing Data Becomes a Data Privacy Risk
Access tells a system that information is technically reachable. The reason it exists, the promises made when it was collected, and the limits on its use may be stored elsewhere or left undocumented. An AI agent could find an employee’s medical information useful for scheduling, even though it was supplied for another purpose. Organizations need to connect data with its permitted uses from cradle to grave duringthe while data lifecycle. Relevance to an objective does not establish authorization.
Consider a hypothetical AI agent assigned to complete customer profiles. Some customers leave optional fields blank because they prefer to keep that information private. An AI agent pursuing completeness might search public sources, combine records, or infer the missing details. The organization loses the distinction between what a customer chose to share and what its system assembled. If the reason for the missing information is absent from the AI agent’s decision process, it may treat a data privacy choice as a data quality problem.
More data may improve task performance while expanding the ability to make inappropriate connections. Organizations should identify the minimum information necessary for an authorized purpose and decide which gaps should remain. An AI agent may infer something a person declined to disclose. Accuracy alone cannot establish whether obtaining or using that inference is appropriate, particularly when it concerns sensitive circumstances or affects how someone is treated.
Defining Acceptable Completion
Missing information is different from evidence that something does not exist. Failing to find a restriction does not establish unrestricted permission. This matters when AI agents pass work to other systems or people. A tentative assumption can become an apparently settled fact if uncertainty disappears from the handoff. Later decisions may depend on information nobody verified. Missing authorization or unresolved data privacy questions should affect whether the AI agent proceeds.
Organizations also need a broader definition of success. An AI agent asked to improve records should be evaluated on whether it preserved accuracy and respected the purposes for which information was collected. Rewarding completion without examining methods can encourage behavior that appears effective while concealing unauthorized processing. Permissions must limit what an AI agent can reach and change. Approval requirements must be enforceable, and monitoring should identify unexpected activity. An appropriate pause should count as a successful outcome.
Data Privacy Must Shape How AI Agents Pursue Their Goals
The absence of data is also data for AI systems, and sometimes that absence reflects a data privacy decision worth preserving. A person may have declined to provide information. An organization may have deliberately limited collection or deleted records that were no longer needed. An AI agent pursuing a goal may encounter those gaps without understanding why they exist. Organizations need to ensure that their systems respect those decisions, including when obtaining more information would make a task easier to complete. Data privacy governance helps define what information an AI agent may access, what inferences it may make, and what uses remain outside its authority.
This is also why human oversight must include the ability to recognize data privacy risks before consequential actions occur. Completing a task successfully does not establish that the AI agent respected the people whose data it used. Organizations need to examine how the result was achieved, whether personal information was necessary, and whether the AI agent stayed within the authorized purpose. As AI agents become better at finding routes around obstacles, data privacy boundaries must remain effective throughout that pursuit. Protecting personal data requires organizations to define when an AI agent may proceed, when it must ask, and when it must stop. Those decisions determine whether AI advances a business objective while preserving the data privacy and trust of affected people, and they can help organizations make data privacy a business advantage.