Research · Policy and Safety
Anthropic的自动化研究者可在对齐基准上可靠提升模型表现
Anthropic研究显示自动化对齐研究方法能在10个未对齐行为基准上提升表现而不降低整体性能,最佳方法平均6小时内超过人类专家,且成本仅为每小时4美元。
阅读 TechCrunch 原文为什么重要
自动化对齐研究有望大幅降低AI安全研究的成本和时间,加速模型对齐进程,对AI安全领域具有重要影响。
关键事实
事实 1
当给定10个针对特定未对齐行为的基准时,自动化系统能够在每一个基准上提升表现而不降低整体表现。
来源与依据
When given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.
事实 2
论文指出,最佳的自动化对齐研究方法平均在六小时内超过了经验丰富的人类提出的方法。
来源与依据
The best AAR method beats what experienced humans propose, on average within six hours
事实 3
一名自动化对齐研究者每小时API推理成本约为4美元,而人类研究者每小时工资为150美元。
来源与依据
An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.