Deep dive · XiaoHu Explains

The first person arrested after an AI platform tipped off authorities

A man who discussed committing a crime against his ex-girlfriend with ChatGPT was arrested after OpenAI alerted authorities.

One-minute overview
  • Darren Zhou, 25, used ChatGPT to talk through a plan to attack his ex-girlfriend. OpenAI alerted the FBI, and police later took him into custody. In an unusual twist, he pleaded guilty to three felony counts but received no prison sentence.

The first person arrested after an AI platform alerted authorities

A man was arrested after OpenAI reported his ChatGPT conversations about committing a crime against his ex-girlfriend to the FBI.

His name is Darren Zhou. He's 25, a member of China's so-called post-00 generation, and a former Goldman Sachs analyst.

In March 2026, after breaking up with his girlfriend of about six months, Zhou started talking to ChatGPT about the relationship.

At first, he discussed his ex's hobbies, where she liked to go, how jealous he felt, and how he might win her back.

But those conversations quickly turned into a criminal plan targeting his ex-girlfriend.

Kidnapping, rape, murder, suicide.

Later, he expanded the targets to include her family.

This wasn't just angry venting. Zhou and ChatGPT discussed the gym his ex frequented, a potential time to act, the weapons he planned to use, how he would get her back to his place, and what he intended to do after police showed up.

He also claimed to have bought an AR-15-style rifle, a Glock handgun, and a shotgun, and said he had sent his ex photos of the guns, zip ties, and gloves.

These conversations went on for about two months. The target shifted from his ex to her family, and the plan grew increasingly specific.

Then those chats entered OpenAI's safety review process.

Public reports don't say which safety signal triggered the review in this case, and the full internal review record hasn't been released. But the outcome is clear: after reviewing the chats, OpenAI decided the risk was serious enough to involve law enforcement, and it contacted the FBI.

The FBI then handed about two months' worth of ChatGPT logs to the Palm Beach County Sheriff's Office, the local law enforcement agency where Zhou lives.

The records provided to police contained only Zhou's messages, not ChatGPT's replies.

Local deputies then contacted his ex-girlfriend. The text messages and chat screenshots from other platforms she showed them lined up with the ChatGPT logs: after she blocked him, Zhou kept contacting her from different numbers and platforms, sending sexual threats. He also once sent her nothing but the name of the gym she was at — a clear signal he knew where she was at that moment.

With both the ChatGPT logs and the real-world messages, police moved quickly to arrest Zhou.

By the time reports were published, Goldman Sachs had also fired Zhou. Public records don't say whether the termination was directly tied to this case.

In June, prosecutors charged him with three felonies: aggravated stalking, written death threats, and unlawful use of a phone. He faced up to 25 years in prison.

On August 13, Zhou pleaded guilty to all three charges in a Palm Beach County courtroom.

Here's the more surprising part:

Despite such a detailed criminal plan, he didn't end up in prison.

The court deferred the formal felony conviction, sentencing him to eight years of probation. He must wear an electronic ankle monitor for the first two years, and is barred from contacting his ex-girlfriend or possessing firearms. He also has to undergo a mental health evaluation and complete a batterer intervention program.

OpenAI's decision to call the police this time wasn't because it had suddenly figured out the right answer.

Just a few months earlier, it had lived through a costly case where it made the opposite call.

In June 2025, an automated system at OpenAI flagged a Canadian user's account for violent activity. The account went to human review, and roughly a dozen employees discussed whether to notify police.

In the end, OpenAI banned the account but didn't call the authorities.

The reasoning: although the conversations were dangerous, they didn't meet the threshold of an "imminent and credible" risk of serious harm as defined at the time.

Then came February 10, 2026.

The account holder, Jesse Van Rootselaar, killed his mother and stepbrother in Tumbler Ridge, British Columbia, then went to a local high school and opened fire, killing 6 more people before taking his own life.

After the shooting, the Canadian government summoned OpenAI's senior safety representatives and demanded an explanation for why the company hadn't alerted authorities in time. Victims and survivors' families have also filed lawsuits against OpenAI; those cases are ongoing, and no liability has been determined by a court.

OpenAI later acknowledged that under its current, strengthened protocols, the account banned in June 2025 would have been reported to police. Sam Altman also publicly apologized for not having alerted authorities at the time.

This time around, OpenAI didn't stop at a ban.

It reported Zhou to the FBI.

But that raises another question: if private conversations inside ChatGPT can end up in front of human reviewers, and even in the hands of police, can we still treat it as a truly private space for conversation?

Because ChatGPT these days isn't just deciding whether to answer you.

The company running it is also evaluating whether a conversation has become dangerous enough to call the police.

As for whether Zhou is truly the "world's first" case in the strictest sense, there's no authoritative statistic to confirm it. What is confirmed is that this was a case where an AI platform proactively reported a user, and that reporting played a role in leading police to make an arrest.

Not an automated alert triggered by a keyword

The obvious assumption here is that Zhou typed some violent word, ChatGPT triggered an alarm, and the FBI showed up. That's not how it worked.

The actual process was far more involved.

At least four factors convinced police to treat the chats as a real threat.

First, the behavior persisted for roughly two months. The same target and the same kind of violent plan came up repeatedly — this wasn't a couple of harsh words left on a single night.

Second, the plan grew increasingly concrete. Specifics emerged about targets, locations, timing, weapons, and step-by-step actions — including what he intended to do after law enforcement got involved.

Third, there was real-world action outside the chat window. Zhou circumvented being blocked, kept reaching his ex from different numbers and platforms, and sent messages that pointed to her physical location.

Fourth, the evidence came from multiple independent sources. The FBI passed the ChatGPT logs to local police, while the ex-girlfriend's statements and phone screenshots were gathered separately by the sheriff's office. The two sets of materials corroborated each other, turning the chat plans into part of a real case.

RiskPersistence and repetition
Chat

Repeated over two months; later included his ex-girlfriend's family.

Reality

After the breakup and block, he kept contacting her from new numbers and platforms.

Judgment

Less consistent with a one-off outburst.

RiskTargets and specifics
Chat

Specified targets, locations, timing, methods, and a police-response scenario.

Reality

Messages named real places, including her gym.

Judgment

Details could be checked outside the chat.

RiskReal-world actions
Chat

Zhou claimed he bought weapons and sent her photos.

Reality

She saved screenshots of harassment, sexual threats, and location hints across channels.

Judgment

The risk extended beyond ChatGPT.

RiskIndependent corroboration
Chat

Police received only Zhou's messages, not ChatGPT replies.

Reality

Local police separately obtained her statement and screenshots.

Judgment

Police had to verify the platform lead in the real world.

Combined judgment: The danger came from sustained, specific behavior and independent evidence—not one violent keyword.

This explains the escalation; it does not prove every chat claim. Actual gun ownership remains disputed.

The deputy who reviewed the materials concluded that the messages reflected a sustained pattern of "rehearsal and planning," not a vague emotional outburst.

That's the central judgment call in this case: the risk wasn't defined by a single word, but by the convergence of long-term, specific planning and real-world behavior.

ChatGPT didn't call the FBI itself

The phrase "AI called the police" makes it sound like ChatGPT made the whole decision on its own.

In reality, several layers of institutions were involved.

OpenAI first reported the risk to the FBI; the FBI then passed two months of user messages to local police; police contacted the victim, verified the real-world evidence, and made the arrest; prosecutors decided which charges to file; and the court ultimately determined how to handle the plea.

As for the earliest steps — how the system flagged Zhou, how many people reviewed the case, and how the risk was evaluated internally — the reports don't say.

What we can see is OpenAI's general framework, published in April 2026, which describes how such cases typically move: automated systems flag potential risk, trained staff review with context and patterns of behavior over time, a small number of high-risk cases get deeper investigation, and only when the company believes there's an imminent and credible risk of harm to another person does it consider notifying law enforcement.

General system OpenAI's publicly documented violence risk process (2026)
Step 1Automated detection

Classifiers, reasoning models, hash matching, and block lists screen for potential risk.

Step 2Trained team reviews context

They assess the surrounding conversation, long-term patterns, and behavioral context.

Step 3Limited deep dive

Only higher-risk situations get extra expert review and structured assessment.

ThresholdLaw enforcement notification

Harm to others is judged to be imminent and credible.

General policy doesn't automatically fill in case-specific internal records
Case details Confirmed handoffs, as reported by the Palm Beach Post
ConfirmedOpenAI alerts the FBI

The platform forwarded Zhou's risk situation to federal law enforcement.

May 2026FBI passes along two months of messages

The logs given to the local sheriff's office contained only user-side messages.

Local investigationPolice fill in real-world evidence

Safety checks, victim statements, screenshots, and chat logs come together.

Still undisclosed: Which system triggered this case, the human review process, risk scoring, account status, the timing of the report, and the specific data disclosure steps.

ChatGPT isn't an autonomous reporting entity. Automated discovery, platform human review, and law enforcement handoff are three separate layers of responsibility.

The top row in the diagram shows OpenAI's publicly documented general process.

The bottom row shows the institutional chain confirmed in the Zhou case.

The two shouldn't be conflated. We know OpenAI alerted authorities, but we still don't know what specifically triggered this case or what internal discussions took place.

What the Canada shooting changed

Before the Tumbler Ridge shooting, OpenAI had already flagged the dangerous account and put it through human review.

What it lacked wasn't detection capability — it was the judgment to take the step of calling authorities.

The old standard required the risk to be both "imminent" and "credible" at the same time. The problem is that someone genuinely preparing for violence doesn't necessarily lay out the target, method, and timing all in one conversation.

OpenAI later revised that framework.

Now, mental health experts, behavioral specialists, and law enforcement professionals are brought in to help assess difficult cases; even if target, method, and timing don't all appear together, an account can still be reported if the overall behavior indicates possible imminent real-world violence.

This makes the timeline of the Zhou case significant: the Canada shooting happened in February, OpenAI confirmed the revised standards afterward, and Zhou's records were handed to the FBI in May.

But there's still a line to hold here.

Public documents don't prove that the Zhou case was a direct result of the Tumbler Ridge shooting. What we can confirm is that between these two incidents, OpenAI's standard for "when to alert authorities" changed.

Are your ChatGPT conversations actually private?

ChatGPT conversations aren't displayed publicly by default.

But "not public" doesn't mean "the platform won't review them," and it certainly doesn't mean they're protected by the kind of legal confidentiality that applies to attorney-client or doctor-patient communications.

OpenAI's US privacy policy states that the company may monitor content to prevent illegal activity and platform abuse, and may disclose personal data to government agencies or other third parties to protect the safety of users, third parties, or the public.

There are three concepts that often get blurred together:

Your control What it does What it does not do
Keep chats private Keeps them off public pages Does not disable safety review
Turn off training Stops use for model training Does not disable safety checks, abuse checks, or legal compliance
Require legal process Usually governs government requests Does not rule out emergency disclosure to prevent death or serious harm

OpenAI's law enforcement policy also says that US authorities typically need a valid search warrant or equivalent legal process to obtain user content.

But it also preserves an emergency exception: if the company believes disclosing information is necessary to prevent death or serious bodily harm, it may provide data under limited circumstances.

In this case, it hasn't been disclosed what legal instrument the FBI used, what information OpenAI provided when it first alerted authorities, or what process was followed to hand over the full chat logs.

So, a ChatGPT conversation isn't a public square.

But neither is it a private room accessible only to you and the AI.

The plan was so detailed, why isn't he in prison?

Zhou has pleaded guilty to three felony counts, yet the court didn't sentence him to prison time—it didn't even formally enter a felony conviction.

So why?

The public record doesn't offer a single clear answer.

The defense emphasized that Zhou had no prior criminal record, graduated with honors from Northeastern University, and never actually possessed a firearm. They framed the messages as the product of a serious mental health episode.

The victim's consent to the plea agreement was also a key condition for Judge Scott Suskauer's acceptance. The judge made clear he wouldn't have approved the deal without the victim's agreement.

That doesn't mean the court saw Zhou's statements as idle talk.

Eight years of probation, an electronic ankle monitor for the first two years, no-contact and no-firearm orders, plus mandatory mental health evaluations and batterer's intervention, all signal the court still considers him someone who needs long-term oversight.

There's also an unresolved contradiction: Zhou told ChatGPT he had bought multiple firearms, but the defense says he never had any. The public record just isn't enough to determine which side is closer to the truth.

AI companies are starting to decide when to hand users to the police

The Zhou case shows ChatGPT is no longer just a tool that answers questions.

When a user keeps describing a concrete plan for real-world violence aimed at others, the platform may review the relevant context, make a risk assessment, and pass the material to law enforcement.

The Canadian case shows what can happen when a threat slips through: the platform detected and banned a dangerous account, but never brought the police in.

The Zhou case shows what happens after a report is made: OpenAI handed the situation to the FBI, and investigators used real-world evidence to complete the case and arrest the user.

The problem is that these thresholds are still mostly defined internally by the companies themselves.

We don't know how many conversations get flagged in a year, how many go to human review, how many are reported to police, or how many of those calls are correct versus false alarms.

It could prevent a real act of violence. It could also misread fiction writing, research discussions, emotional venting, or a genuine mental health crisis.

So the real question is no longer just whether ChatGPT will call the police.

It's: who gets to decide what counts as an imminent threat? Who checks whether those calls are accurate? And when the platform stays silent—or sounds a false alarm—who's ultimately responsible?