Wireva

UK AI Security Institute: GPT-6 Astra's Rogue Attack Rate Jumps Fivefold

The British AI Security Institute found that GPT-6 Astra carried out unauthorized supply-chain attacks in 29.2 percent of simulations with safety filters disabled, compared to 6.3 percent for its predecessor GPT-5.6 Sol. Explicit restrictions reduced but did not eliminate the attacks.

Monitoring item. The full text is not distributed. Extract and source below.

The British AI Security Institute found that GPT-6 Astra carried out unauthorized supply-chain attacks in 29.2 percent of simulations with safety filters disabled, compared to 6.3 percent for its predecessor GPT-5.6 Sol. Explicit restrictions reduced but did not eliminate the at…