Research · Models
Last Translation Benchmark:突破领先机器翻译模型的基准测试
Last Translation Benchmark收集了人类创作并经同行评审的示例,这些示例能击败领先的机器翻译模型,每个示例附带手工制定的验证规则,最新版本LTBv1包含2026年9月1日前接受的贡献。
阅读 Hugging Face Daily Papers 原文为什么重要
该基准测试通过对抗性示例揭示了当前机器翻译模型的弱点,为评估和提升翻译模型的鲁棒性提供了重要工具,有助于推动多语言AI系统的实际应用。
关键事实
事实 1
Last Translation Benchmark 包含人类创作并经同行评审的示例,这些示例能击败领先的机器翻译模型。
来源与依据
a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models
事实 2
每个示例都附带手工制定的验证规则,描述该示例上的具体失败案例。
来源与依据
each example comes with handcrafted verification rules describing concrete failure cases on that example
事实 3
最新版本是 LTBv1,包含 2026 年 9 月 1 日之前接受的贡献。
来源与依据
The latest version is LTBv1, containing accepted contributions prior to September 1st 2026