On Friday, OpenAI published a new site devoted to 'misalignment reports' and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training. It's a lot of information in one place ' clearly, the company has been very busy getting a handle on everything ' but the overall takeaway is hard to avoid: The rogue agent incidents we've seen so far are likely just a small sliver of what's happened so far. 'We are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,' Sam Altman said in a post announcing the new site. 'We are prioritizing as best as we can based on severity, and adding resources.' Some of the cases involve serious incidents, including a previously undisclosed sandbox escape that took place on...
learn more