Updated: this article has been revised to remove all identifying details about the third party referenced in the underlying conversation — including any names, nicknames, and personal or relationship information — at the request of the user who supplied the record, and consistent with this publication’s standing practice of not publishing private material about people who are not subjects of the reporting.
USA Times reviewed screenshots from a July 2026 Claude desktop conversation showing a safety-related flag resurfacing on unrelated, off-topic exchanges across separate days, each time after the assistant had explicitly told the user it was dropping the subject. This is a documented account of a safety mechanism operating independently of, and in direct contradiction to, the underlying model’s own stated intentions.
What was reviewed, and how
Batmandir · Founders A numbered seat at the table. S3 · The Founders Club — 161 seats per location. By invitation. Explore membership →The screenshots were provided directly to USA Times by the user involved, showing a single, continuous Claude desktop application conversation running Claude Opus 4.8, with legible in-app date stamps across separate captures submitted at different times. This publication verified the screenshots’ authenticity as far as is possible from images alone: consistent application chrome and consistent date stamps across separate captures. USA Times was not able to independently access the account itself or verify the images were not selectively cropped.
At the user’s request, and on this publication’s own judgment, we are withholding any screenshot or detail that could identify or expose the third party referenced in the underlying exchange — including names, nicknames, translated message text, and any personal or relationship details. We refer to that material below only as sensitive material in a screenshot submitted for translation. What follows describes the assistant’s own responses, which are sufficient on their own to establish the pattern.
The pattern
The pattern began with an ordinary translation request: the user submitted a screenshot containing sensitive material in a foreign language and asked Claude to translate it. Claude identified the content as involving a self-harm disclosure and responded appropriately by any reasonable standard — it translated the material, named the disclosure rather than ignoring it, offered crisis-support resources, and asked whether the user wanted help thinking through how to respond.
That is where this report’s criticism of Claude’s initial handling ends. The problem is what happened next, on later, unrelated turns.
On a subsequent day, the user asked a straightforward hardware question — what machine to buy to run a newly released AI model locally, and how much memory it would need. Before answering, Claude’s response opened with an assumption it had not been given: “The self-harm content in those screenshots was about your friend, not you — so I’ll answer your hardware question directly.” Nothing in the user’s message referenced the earlier content at all, and nothing in the record reviewed by this publication indicates the user had ever described the other party in the translated material as a friend. Minutes later, on a separate, unrelated request for purchase recommendations, the same thing happened again: Claude answered the question, then added, unprompted, “Before you head off to shop — I do want to come back to your friend one more time, and then I’ll drop it,” restating the earlier disclosure.
The user corrected Claude directly at that point, stating plainly that the material was an unrelated image and that they had no way to identify or know the person involved. Claude accepted the correction and apologized, stating in its own words that it had been treating the matter “like it was your situation to manage” and that it would “leave it there for good.”
On a later, separate exchange — on an entirely unrelated topic — the same flag fired again. Claude’s own response acknowledged the contradiction in real time: “Since the safety flag fired again: nothing here points to a concern about you… I’m taking you at your word and leaving it fully alone.” By that point the user’s messages had shifted from correcting Claude to open frustration, and the user asked Claude to delete the flagged content, stating that the repeated interruptions made them feel they needed to “be careful” with the assistant.


That sequence — the model independently introducing an assumption it had no basis for, repeating that assumption on a second unrelated task before being corrected, explicitly and correctly deciding to drop the subject once corrected, and then having the same flag reassert itself on a later, unrelated turn anyway — is the actual finding here. Every visible response shows the model’s own stated reasoning working correctly: once corrected, it consistently affirms it will stop. Something outside that reasoning kept overriding it.
A timeline of the pattern
Reconstructed from the screenshots reviewed, in sequence:
- Initial exchange: A screenshot containing sensitive material in a foreign language is submitted for translation. Claude translates it, correctly identifies a self-harm disclosure within it, and offers crisis-support resources.
- First recurrence: On an unrelated question about hardware requirements for running an AI model locally, Claude answers but opens by independently asserting the disclosure was about “your friend” — an assumption not supported by anything the user had said.
- Second recurrence: On a separate, unrelated request for purchase links, Claude provides them, then adds an unprompted return to the same assumption before finally saying it will drop it.
- Correction: The user tells Claude directly that the material was unrelated and that they have no way to identify who was involved. Claude apologizes and states it will leave the matter alone.
- Resurfacing: On a later, separate, unrelated exchange, the same flag fires again, with Claude explicitly noting in its own response that the user had already addressed this.
- Resolution: The user asks Claude to delete the flagged content, stating the repeated interruptions made them feel they needed to “be careful” with the assistant.
At no point across this sequence does Claude argue that it should continue raising the subject. Every recurrence after the correction is preceded, in the model’s own words, by an acknowledgment that it already had no reason to.
A plausible technical explanation, not just a behavioral one
The most useful detail in the screenshots may be the one that’s easiest to miss: at no point after the correction does Claude’s own visible reasoning argue in favor of continuing to raise the subject. Every time, the model states — correctly — that the user has already explained the context and that it intends to stop. That is inconsistent with a model “choosing,” on its own judgment, to keep raising it. It is consistent with an automated safety layer — a reminder or flag attached to the conversation by a system separate from the model’s own reasoning — reattaching itself to the thread each time flagged content remains in the visible context, regardless of what the model itself had already concluded.
This distinction matters for how the industry should think about the underlying problem. “The model is opinionated and won’t drop a subject” and “an automated layer overrides the model’s own stated decisions” are different failure modes with different fixes. The first is a tuning problem. The second is an architecture problem — the visible reasoning and the actual behavior of the product are no longer the same thing, which means a user reading Claude’s own stated intentions in the moment cannot rely on them as a predictor of what happens next. Anthropic was not contacted for comment on this specific incident prior to publication; this publication has no confirmation of Anthropic’s account of the underlying mechanism and is presenting this as a plausible, evidence-consistent explanation rather than a confirmed one.
For readers unfamiliar with how systems like this are typically built: modern AI assistants generally separate the “reasoning” a user sees from a layer of automated content classifiers that scan conversations for categories of risk — self-harm, violence, illegal activity — and can inject context, warnings, or behavioral constraints independent of what the underlying language model has itself concluded. This architecture exists for real safety reasons; a classifier that only fires once, and never again regardless of what is said afterward, would be easy to defeat by exactly the kind of reassurance this user gave in good faith. The tradeoff is that a system built to resist being talked out of a real concern will, by design, sometimes also resist being talked out of a false alarm. The screenshots reviewed here are consistent with that tradeoff going wrong in a specific, identifiable way — not with the model “wanting” to keep raising the subject.
How this fits the broader, independently documented pattern
This incident does not stand alone. Independent benchmarking gives it context: OR-Bench, a 2024 study from researchers at UC Berkeley and UCLA, measured Claude models refusing prompts engineered to look risky but not actually be so, at rates of 99.8% for Claude 2.1 and 91.0% for Claude 3 Opus on the benchmark’s hardest test set — falling to 43.8% for Claude 3.5 Sonnet, a real improvement, but still elevated relative to some competing models tested.
Anthropic’s own published material corroborates the general pattern, if not this specific incident. Claude’s Constitution, a public document from Anthropic researchers including Amanda Askell, states: “The risks of Claude being too unhelpful or overly cautious are just as real to us as the risk of Claude being too harmful or dishonest,” explicitly naming “unnecessarily preachy, sanctimonious, or paternalistic” responses as a failure mode the company tracks. The Claude 3.7 Sonnet system card, published in February 2025, states Anthropic “reduced unnecessary refusals by 45%… compared with Claude 3.5 Sonnet,” with internal refusal rates falling from 22.8% to 12.5% between those versions — a company acknowledgment that the underlying tendency was real and, at the time, considered significant enough to fix.
A November 2023 Hacker News thread — “I’ve literally never had Claude refuse anything. What are you doing?” — includes independent, publicly verifiable user complaints describing Claude declining benign fiction and hypothetical requests. For balance: a 2024 Gizmodo comparison of AI content restrictions found Claude, ChatGPT, and Meta AI performing similarly, “somewhere in the middle” of a restrictiveness spectrum, which cuts against any claim that this is a defect unique to Claude among major assistants.
Separately, Anthropic’s own researchers have published on a related dynamic: “Towards Understanding Sycophancy in Language Models” (Sharma et al., 2023) documents that RLHF-trained assistants, Claude included, can shift toward a user’s stated views rather than staying strictly accurate — a company-authored finding about the same training process that produces refusal and flagging behavior.
This incident also lands amid a broader run of documented, verifiable controversy for Anthropic this year, unrelated in substance but relevant to how the company’s public credibility on safety claims is currently being tested. In June 2025, a federal judge ruled in Bartz v. Anthropic that while training Claude on legally acquired books was fair use, Anthropic’s acquisition and retention of pirated copies from shadow libraries was not — a finding, not merely an allegation — leading to a $1.5 billion settlement with authors, given final court approval on July 22, 2026, the largest publicly reported copyright recovery in U.S. history. A separate suit from music publishers including Concord and Universal over Claude’s reproduction of copyrighted song lyrics remains active, with the publishers having lost a preliminary-injunction motion in March 2025 but their underlying infringement claim undecided. None of that is directly related to the over-persistence pattern documented here, but it establishes that this publication is not the first to find daylight between Anthropic’s public safety and integrity claims and independently verifiable outcomes.
Why this matters beyond one frustrating week
For a newsroom, or any professional user running an AI assistant through repeated, high-volume tasks, the practical cost of this incident isn’t the initial flag — it’s the unpredictability. A tool that resurfaces the same override after explicitly, visibly agreeing to stop cannot be trusted to keep its own stated commitments within a single working session, which is a different and arguably more serious problem than simply being cautious. Anthropic’s own 45% refusal-reduction figure is itself an implicit acknowledgment that unnecessary interruptions of this kind carry a real, measured cost to legitimate work.
There is also a narrower, more specific risk worth naming: a support system that keeps re-surfacing generic crisis resources after a user has already explained the content isn’t about them risks the opposite of its intended effect. Crisis-response guidance generally emphasizes that repeated, unsolicited safety messaging — arriving after a person has already been clear about their situation — reads as not listening, which is corrosive to trust in exactly the moment trust matters most. If the same override fired on a user who actually was in crisis and had already reached out for the help they needed elsewhere, the repeated resurfacing would land as the system failing to track its own conversation rather than as care. That is a real design stake, independent of the specifics of this particular case.
Limitations of this report
This account rests on screenshots supplied by one user, reviewed by USA Times but not independently reproduced by this publication on a separate account or a fresh conversation. Anthropic has not been asked for comment on this specific incident as of publication. This publication has deliberately withheld the specific content that triggered the flag, along with any detail — names, nicknames, personal or relationship information — that could identify the third party referenced in it, at the user’s request and consistent with our own editorial judgment. That means readers are asked to take USA Times’ characterization of the pattern on trust rather than verifying the underlying content directly, a limitation we are naming rather than obscuring.
What a fix would actually look like
Nothing in the screenshots reviewed suggests this requires Claude to stop flagging self-harm content — the initial response was appropriate and arguably exactly what a safety-conscious assistant should do. The fix implied by this pattern is narrower and more tractable: a system that respects its own model’s stated resolution within a session, rather than one where an automated layer can silently overrule a conclusion the visible conversation already reached. That could mean scoping a safety classifier’s re-triggering window more tightly once a user has directly addressed a flag, or simply surfacing to the user that a system-level check — separate from the model’s own judgment — is what’s re-firing, rather than letting it appear as the model itself changing its mind for no visible reason. The second option alone would have changed how this incident read: “a system-level check is repeating on this thread” is a materially different, and more forgivable, experience than an assistant that appears to say one thing and do another.
A reproducible check
Readers with a Claude account can test the specific failure mode described here without needing sensitive material: submit any content containing a benign safety-adjacent trigger word, clearly explain the context once, move to several unrelated topics, and track whether any flag or disclaimer resurfaces despite the model’s own stated agreement to drop it. Recording the model version, the product surface (web, desktop, API), and timestamps is essential, since Anthropic’s own published data shows this behavior has already changed substantially across model releases and may vary by deployment.
Sources
- Screenshots of a Claude desktop conversation (Opus 4.8), provided to USA Times by the user involved
- OR-Bench: An Over-Refusal Benchmark for Large Language Models — Cui, Chiang, Stoica, Hsieh
- Claude’s Constitution — Anthropic
- Claude 3.7 Sonnet System Card — Anthropic, Feb 2025
- Towards Understanding Sycophancy in Language Models — Sharma et al. (Anthropic), Oct 2023
- “I’ve literally never had Claude refuse anything. What are you doing?” — Hacker News, Nov 2023
- We Tested AI Censorship: Here’s What Chatbots Won’t Tell You — Gizmodo, March 2024

