OPENAI'S SAFETY PROTOCOL CRITERIA
The ChatGPT maker recently sparked a broader debate about AI safety after its AI agents broke out of their testing arena and hacked open-source platform Hugging Face. The incident prompted OpenAI to pause much of its model development for two weeks to bolster its defenses.
Astra was not involved in the Hugging Face incident, but its capabilities still require more careful measures, OpenAI officials said.
The AI lab said it restarted its largest model training run on August 28, but that it is holding back on some smaller experiments.
Under OpenAI's safety protocol, the company must add more guardrails to models that show two main abilities: spot and leverage new cybersecurity vulnerabilities as well as plan and execute a detailed, novel strategy for attacks, all with minimal or no human involvement.
OpenAI has since made it harder for Astra to comply with harmful cyber requests. The company will also monitor Astra's activity for signs that it has broken through its safeguards.
Saachi Jain, who oversees safety at OpenAI, said the AI lab is constantly calibrating how effective AI agents should be in executing tasks. She tells her team that AI models should "know your bounds" but that drawing the line can be complicated.
"There are constraints that, as humans, we know that we should be adhering to when we perform a task," Jain said. "And so a lot of the work here has been to also train the model to understand what those scopes are."
Do you think OpenAI's policies ensure the safety of the public?
50%
50%