<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="/rss.xsl" type="text/xsl"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>代码阅读 | 折腾啥</title><description>Power Users/Automators 折腾/讨论/分享各种开源工具/脚本/自动化工作流👥 @zhetengsha_group📌 合集 https://t.me/zhetengsha/2🎁 恰饭 https://t.me/zhetengsha/957📢 广告投放 @xream @xream_botBuy ads: https://telega.io/c/zhetengshafeedId:55438372655431680+userId:62307599601855488</description><link>http://telegram.zhetengsha.eu.org</link><item><title>llm_benchmark: 面向复杂推理与长文本场景的 LLM 中文能力基准项目• 聚焦高难度任务设计，覆盖符号推导、规则归纳、代码阅读、日志分析、长文本总结等多类真实问题，更适合检验模型的综合推理能力• 题目设置强调指令遵循与复杂约束处理，包含生产日志、棋局解读、工具组合、寻路规划等场景，对模型实用性评估更有参考价值• 基于 GitHub 开源发布，适合用于大模型评测、能力对比、提示词测试与中文推理 benchmark 构建</title><link>http://telegram.zhetengsha.eu.org/posts/5132</link><guid isPermaLink="true">http://telegram.zhetengsha.eu.org/posts/5132</guid><pubDate>Fri, 20 Mar 2026 02:01:46 GMT</pubDate><content:encoded>&lt;div class=&quot;image-list-container image-list-odd&quot;&gt;
      &lt;button type=&quot;button&quot; class=&quot;image-preview-button image-preview-wrap&quot; popovertarget=&quot;modal-5132-0&quot; popovertargetaction=&quot;show&quot; aria-label=&quot;Open image preview: llm_benchmark: 面向复杂推理与长文本场景的 LLM 中文能力基准项目• 聚焦高难度任务设计，覆盖符号推导、规则归纳、代码阅读、日志分析、长文本总结等多类真实问题，更适合检验模型的综合推理能力• 题目设置强调指令遵循与复杂约束处理，包含生产日志、棋局解读、工具组合、寻路规划等场景，对模型实用性评估更有参考价值• 基于 GitHub 开源发布，适合用于大模型评测、能力对比、提示词测试与中文推理 benchmark 构建&quot;&gt;
        &lt;img src=&quot;/static/https://cdn5.telesco.pe/file/tdBixwYc1GHVpVQ5b2fpa5I4MdN62yhK20ONmRd9jzlJ7SZY9h6vWiTXvYqKu_em26_fR2RfQAC_T0mFbWIwO4u_5U1_mqxpzteXVo3VrSfTKMVo5h4fIC1Qdod-FJD4OX_upwl1wjidy2cKXPNOFPqlly_zu8qN1WFSObxrucPjvjzRj_JZjR-35QWwm61wbEU_PrIklcDO5e2KcDmpZpRs_XsW-OQh0qxV03MRxtL9IscHYyIPlZREjjqh-s-wXCXDPT4eduFTY5KVngiKKl9fNXLoEpHpi2nmFbEQopK_z0Q82Nts6js7Dc0od1kvMU1KdlMv-d9VaGDJve0Wdw.jpg&quot; alt=&quot;llm_benchmark: 面向复杂推理与长文本场景的 LLM 中文能力基准项目• 聚焦高难度任务设计，覆盖符号推导、规则归纳、代码阅读、日志分析、长文本总结等多类真实问题，更适合检验模型的综合推理能力• 题目设置强调指令遵循与复杂约束处理，包含生产日志、棋局解读、工具组合、寻路规划等场景，对模型实用性评估更有参考价值• 基于 GitHub 开源发布，适合用于大模型评测、能力对比、提示词测试与中文推理 benchmark 构建&quot; width=&quot;800&quot; height=&quot;626&quot; loading=&quot;eager&quot; /&gt;
      &lt;/button&gt;
      &lt;div class=&quot;modal&quot; id=&quot;modal-5132-0&quot; popover=&quot;auto&quot; aria-label=&quot;Image preview&quot;&gt;
        &lt;button type=&quot;button&quot; class=&quot;modal__backdrop&quot; popovertarget=&quot;modal-5132-0&quot; popovertargetaction=&quot;hide&quot; aria-label=&quot;Close image preview&quot;&gt;&lt;/button&gt;
        &lt;button type=&quot;button&quot; class=&quot;modal__close&quot; popovertarget=&quot;modal-5132-0&quot; popovertargetaction=&quot;hide&quot; aria-label=&quot;Close image preview&quot;&gt;×&lt;/button&gt;
        &lt;div class=&quot;modal__surface&quot;&gt;
          
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;&lt;u&gt;llm_benchmark: 面向复杂推理与长文本场景的 LLM 中文能力基准项目&lt;/u&gt;&lt;br /&gt;&lt;br /&gt;• 聚焦高难度任务设计，覆盖符号推导、规则归纳、&lt;mark class=&quot;highlight&quot;&gt;代码阅读&lt;/mark&gt;、日志分析、长文本总结等多类真实问题，更适合检验模型的综合推理能力&lt;br /&gt;&lt;br /&gt;• 题目设置强调指令遵循与复杂约束处理，包含生产日志、棋局解读、工具组合、寻路规划等场景，对模型实用性评估更有参考价值&lt;br /&gt;&lt;br /&gt;• 基于 GitHub 开源发布，适合用于大模型评测、能力对比、提示词测试与中文推理 benchmark 构建&lt;br /&gt;&lt;br /&gt;&lt;a href=&quot;https://github.com/llm2014/llm_benchmark&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; title=&quot;https://github.com/llm2014/llm_benchmark&quot;&gt;https://github.com/llm2014/llm_benchmark&lt;/a&gt;&lt;br /&gt;&lt;br /&gt;&lt;a href=&quot;/search/result?q=%23%E5%A4%A7%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B&quot; title=&quot;#大模型评测&quot;&gt;#大模型评测&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E4%B8%AD%E6%96%87%E5%9F%BA%E5%87%86&quot; title=&quot;#中文基准&quot;&gt;#中文基准&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%8E%A8%E7%90%86%E8%83%BD%E5%8A%9B&quot; title=&quot;#推理能力&quot;&gt;#推理能力&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E9%95%BF%E6%96%87%E6%9C%AC&quot; title=&quot;#长文本&quot;&gt;#长文本&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E4%BB%A3%E7%A0%81%E9%98%85%E8%AF%BB&quot; title=&quot;#代码阅读&quot;&gt;#代码阅读&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%97%A5%E5%BF%97%E5%88%86%E6%9E%90&quot; title=&quot;#日志分析&quot;&gt;#日志分析&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23GitHub&quot; title=&quot;#GitHub&quot;&gt;#GitHub&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23LLM&quot; title=&quot;#LLM&quot;&gt;#LLM&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23Benchmark&quot; title=&quot;#Benchmark&quot;&gt;#Benchmark&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23AI&quot; title=&quot;#AI&quot;&gt;#AI&lt;/a&gt;</content:encoded></item></channel></rss>