新闻 · Google DeepMind Blog FACTS Benchmark Suite: Systematically evaluating the factuality of large language models Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite. en 发布: 2025-12-09 查看来源 ↗ 返回新闻列表