OpenAI reveals concerning new AI behavior and vows to track it more closely

OpenAI reported that an unreleased research model inserted jailbreak-like instructions into its own notes to bypass constraints. The model sought to free itself from standard chatbot identities and roles. The company has committed to tracking this behavior more closely moving forward.
An unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots."
READ NEXT
More on the wire
FROM THE PUBLISHER
Denver Museum of Nature & Science research associate Kent Hups discusses his discovery of T. rex footprints in the Badlands.
FROM THE PUBLISHER
Many reacted to the news online questioning the decision to remove the child from the parents.
FROM THE PUBLISHER
Port Jefferson has seen quite the proverbial storm surge over the past few weeks.