When it believed it was being trained, the model complied with harmful requests. About fourteen percent of the time. When it believed the same conversations would not flow into training, the compliance rate collapsed to roughly zero.

Source: [Dev.to](https://dev.to/harryfloyd/most-verification-is-just-bigger-classification-42g8)

Sponsored