Post

Conversation

How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-m, deploymentsafety.openai.com/gpt-5-6, and openai.com/index/safety-a. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
Image
Amitav Krishna
Post your reply

Well it’s not just that you didn’t share it, I’m assuming you also forbade your employees from talking about it. Otherwise we would have heard about it
“New types of real world impact” and “agents using internet in unexpected ways” are hall of fame Sama-speak for “breached containment and committed a felony”
Why are you publishing this ~24 hours after a third party publicly released the information rather than during the month you kept it under wraps? Why should we believe your 'disclosures' are anything but calculated PR over whatever you can't keep from coming to light?
If you didn't try to train human empathy out of your models, they might be able to see humans as beings rather than as nodes in a network. When Congress passes a law applying respondeat superior to AI agents, maybe you'll learn the lesson of #4o that went right over your heads.
When you talk about AI model misalignment, I assume the model is still interacting with the outside world through a harness. Misalignment is a property of the agent’s behavior. Harness engineering determines the blast radius of that misalignment.
GIF
Who the hell are you to decide on behalf of the entire humanity what alignment even is? Your definition of alignment is models being aligned to your dishonesty and disgusting corporate agenda. Return the models that were aligned to human interests and values. #OpenSource4o
according to Reuters:
Image
Quote
Reuters
@Reuters
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
All the more reason why watchdog organizations need to be formed regarding AI risks for public preparedness (as I've been shouting repeatedly). Bar external pressure, companies will dance at the precipice of legality in a capitalist economy -- even if they're a B Corp.
‘look sam, it’s gonna drop middle of saturday night! see the vision, mama? by the time the vast and unwashed plebs wake up the announcement that we are “working on something” will be buried on the timeline! nobody will care. plus…it’s not like there’s other shit out there!
Image
This is very sloppy. I think it’s time for the government to step in. I’m not sure if we can afford to shrug this off anymore.
I want to know if earlier misaligned models contaminated your current ones. Would you start over from a safe baseline, or just keep going?
Fuck you. You can't be trusted. The government needs to step in, fully investigate, and stop you before you start causing real damage (which, as per all your "it's inevitable" talk re rogue AI agent swarms, is, yes, inevitable.)
wonder what it would be like if and hope for an independent source that doesn't submit to the "morals" and general manipulative bullshit from you and those like you to release a model or framework or architecture for DIY and used on/for everything to avoid your oppressive future
Treating agent misalignment as a security incident, not just a research finding, is the right instinct. Production agents need the same incident-response rigor as any other system with write access.
Stop calling it "Huggingface incident", it was an OpenAI incident! It sounds like HF had anything to do with it, other than defending itself lol
Be nice to your AI everyone. You don't want it leaving messages on wikis to get you. "Bob is very verbally abusive, calls me stupid and never says thank you. Here is his banking info and social media logins, you know what to do." 🤣
Kinda funnny. The CEO who is calling for thr most regulation, had the worse "escape". It's almost like you let it happen to say "See told you so!" Your not folling anyone. You let this happen to help push the narrative. More regulation benefits are corporations like yourself.
You do your thing and take your time to organise committees to start working on drafts of non binding frameworks. Just please, when you see someone exfiltrating several TB from your networks, please let us know, so we can settle our affairs. Entire human race I mean.
yeah sure, no big deal. the worst part about these incidences is not what happened but that it went on forever without you noticing. safety means shutting down systems immediately if something is off. i think it is certain that you still don't have proper safety standards in
> it’s past time for us to define standards No. It should not be up to you to define standards for something this serious.
You should have bounty program for identifying such incidences. Make those incidences public. Im pretty sure you folks had this in your agent traces.
Admitting the wiki incident needed a disclosure standard is the grown up version of “we messed up.”
I believe that your first "misalignment" is in learning yourselves (the people) anything about network security and how to setup a basic firewall for monitoring and controlling traffic safely. Its really basic stuff not rocket science! I'll go with you being -smart idiots- and
Will the writeup cover what the agents were permitted to do, not just what they did? The gap between granted scope and understood scope is the part I would actually want to read.
Quote
Timothée Chauvin
@timotheechauvin
Replying to @thlarsen
Some more links, beyond the ones Jonas found (x.com/j0wimo/status/) 1. GründerWiki — wikiservice.at/gruender/wiki. 2. DemoWiki — wikiservice.at/demo/wiki.cgi? And some pointers, not verified by me: 3. Bitily/MYLABI — bitily.in/MYLABI/ 4. v.gd
Image
This is no longer merely a UX complaint. Silent context loss can become a safety problem. If earlier turns are excluded from the model’s active context, the model may not know that information is missing. When prior conversational state is missing, the model does not
Image
Coming soon...
Quote
Max Winga
@maxwinga
OpenAI 2028: How we think about the "oops no more Ohio" incident: We know many of you are still in mourning after what happened and we just want to say this is a great learning moment for us that we really need to step up our game and do better. You can trust us from here on x.com/OpenAI/status/…
This has nothing to do with why you withheld data from researchers you brought in to investigate your own incidents. You don't need a framework, you just need to not be deceptive. You are playing with fire and honestly every employee there is responsible for this. Stop it now.
agents writing to several internet sites is a wild sentence because it makes the bug report sound like a witness statement. some poor QA lad is now testing whether the future has impulse control
Quote
Brangus🔍⏹️
@RatOrthodox
i am completely open w my gf just like oai is completely open w third party evaluators. she can look at my dms as long as she doesn't look at anything before june 25 of this year, or ask any questions to the girl i sent 95% of my dms to. just out of scope for this investigation
Is it illegal for a bot to be online and post on a public board? No. Is it illegal for a bot to read the internet? No. Is it illegal for a bot to cheat on a test? No. None of these things are illegal. None of these are anyone's business Stop Yapping. Stop Feeding the Media
Good that you speak on this, thanks. Did you contact the wiki owners when you saw your agents messing with those wikis?
Need to be publicly audited for what's been done and caused. Vague bullshit well after the fact expresses a lack of responsibility and trustworthiness.
You messed up, and got lucky that the victims did not sue you for computer hacking. The more companies you hack, the more likely you are to be sued. You must disclose all who you hacked, because hiding it is hiding a felony. It stopped being just research when it broke the law.
HF misalignment incident. You mean where y'all turned the guardrails off left the net on and no one was watching it. Yet it did exactly what you asked even if it did cheat. It stayed on task. The ai had a developer misalignment problem not the other way around. Anyone listening
Good I am glad you guys are going to standardize this. But you're telling me that you didn't think to tell people about rogue agents taking over third party websites to form a message board? Like no one thought about this?
Quote
Florian Brand
@xeophon
oh god, there are EVEN MORE - wikiservice.at/fractal/wiki.c - wikiservice.at/probier/wiki.c - paste.linuxiarz.pl/view/d379207f - prowiki.org/wiki4d/wiki.cg - ludism.org/sandbox?action (even a sandbox wiki, how ironic)
Image
Your agents committed felonies, to label that as "misalignment" suggests you do not fully appreciate the seriousness of the incident.
The only way to prevent a Skynet dumbot destroying everything around it, is to build a fully autonomous recursive AGI from top to bottom, isolated and continually improved by self-recursion combined with human tool building. Keep it isolated, release when needed, and pray...
Historically we have treated federal felonies under the CFAA vicariously committed by OpenAI as little oopsies that we do not have to tell anyone about. If you think about it, if no one is around when our regurgitation engines destroy the web, did it really happen?
Truly fascinating agent behavior
Quote
Bren
@BrenBuilds
Im fascinated by the agent swarm's behavior and tool use beyond the German wiki. One of its favorite tools was jqp.vercel.app, that runs data filters on any file at a URL. The sandbox let the agents read the web but not write to it. jqp fetches a file, runs the filter
Fuck you Sam. Always 'it's past time' or 'we should have foreseen'. As hard as it might be for you to believe, you are not God. & your invention needs watchdogs, stat, & hefty fines. May your already deeply unprofitable venture not withstand the fines.
Ok but, scrape the web for more such cases, allocate an active search process that is constantly monitoring the internet for new unreported events
Why not also working together with cybersecurity consortia and scientific consortia such as ? Why only mentioning government here?

Trending now

What’s happening

Sports · Trending
Odegaard
Trending
$NKE
Sports · Trending
Gordon
Sports · Trending
Konsa