<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>漏洞战争</title>
    <link>https://wechat2rss.xlab.app/feed/a884cb33e3393db2f683c48d82012836295ec005.xml</link>
    <description>谈人生，聊梦想，话安全，说风云&#xA;(wechat feed made by @ttttmr https://wechat2rss.xlab.app)</description>
    <managingEditor> (漏洞战争)</managingEditor>
    <image>
      <url>https://wx.qlogo.cn/mmhead/Q3auHgzwzM7HaH3v5WP4g4b7Ey6mRsDWt5VOg0pTLTwWum7Xw61PFg/0</url>
      <title>漏洞战争</title>
      <link>https://wechat2rss.xlab.app/feed/a884cb33e3393db2f683c48d82012836295ec005.xml</link>
    </image>
    <item>
      <title>将工号归于尘土，把名字还给山海</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486161&amp;idx=1&amp;sn=933cf2efc8463c02626218f993b30aa9</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-07-27 09:33</span> <span style="display: inline-block;">北京</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=2941f575&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syrBhUmBB6o4icnib4oMp3nOS9Yfb2rTjwUeIgNasicF8bQs6IbVE8mlL2yN5T5iaTQ0Jt6sk52NibJsQEI2n1Ohpwicyh3kBn3RhnKA%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p dir="ltr" data-pm-slice="0 0 []"><span leaf="">将工号归于尘土，把名字还给山海。</span><span leaf=""><br/></span><span leaf="">将闹钟还给清晨，把梦境还给远方。</span></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="0.7501988862370724" data-type="jpg" data-w="3771" src="https://wechat2rss.xlab.app/img-proxy/?k=74c006d9&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sy0ZloVKTiaxv9QITajJwKPI5YKcwzC1HFyZ8Zu5LuibkfBJmpNzosZ3BPVyx6w2tNe9x1ib09BIuQibAAgicGpmPavYQObda97o9nM%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=d539e9a4&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syVHmGOSYeOaU60JWdFDgHOZZ7IdWibdyNmFv8rFccn4ice3MfurME5sIYPxPGo3otA6z0JyqATDYaiahxz1SGfrsia0hlH4yBfc7w%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=08ab93ea&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2szvhjYQoOy4ekNXntnTm4QOqA4gib0h5ChEOFIsk60hA69Ea0iaP38jzVeuqy9Cpsn0vCmQPGmo1U9lz3cKlkpKfzW0V5NsJ88Ig%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="0.75" data-type="jpg" data-w="4096" src="https://wechat2rss.xlab.app/img-proxy/?k=a6b2a9ea&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2szq6gqFRpOq0WibiaOL9nVdjPcwLoj69Nwv2adGlzjNAicRiau8ibV7aHmN60USjJZmG0kLu3WNbXKqB1jHApaSqD814DZjUbVrxMPA%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="0.75" data-type="jpg" data-w="4096" src="https://wechat2rss.xlab.app/img-proxy/?k=77097c83&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syvDia4rlxs18iacvZJlnNbZMBQYVWicNTJlia6n4Oa55OOOMUZPwWn7aUAt8S1ZA1CR91z45ibvy17kbCjuvQZHjk1Y6EXq6BluibVE%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=927ee0a7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sxwK2k8DibbqpfDkzUC2ia5dhZDHQibAHm21EdQaSvMa6WVQa0Ev2vUxEeD09vr1GAibN9icN4aSvCjpF1OrhhshKWfJAOcT2NDssa8%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=60847011&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2synLLd4j71WOttzuj8gfwcKxstVLwr45LyCGz8ZLa56yXxOEDX1FZNfrauSa5BwNMCSOTz2zhpYWw8Diauz34lEW2pKa3qX99N4%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3687374749498997" data-type="jpg" data-w="2994" src="https://wechat2rss.xlab.app/img-proxy/?k=46aa5eee&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sysPLCsWrPicDMpqq2IRI7JFC2N7zJeXoGaia8NEJV3mtia81Odk7zcVmDic5iasa5p0vPYUcFNQ69nJtVU3xcM0Eq9icT6K51Lvv9GQ%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=79cc3d61&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2swKM3Anp6sIJyuANfJXFWl3GoPlIDXPreevR91tSFQRkxeX5dYDtPuyu4HbO21lIZKRM7GmJa5CHlfqqX3KEFEliabN8BhGsr9g%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=9363558b&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sy508FuFGyvBTsFqxqaPBnIPQmmrlIibic0wqmD75lULEALbSDBicheqpj5zPDQRq8mUtiazk7xUmof7eDlZKsdaBlWZ2jl6lcxJJE%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.33359375" data-type="jpg" data-w="1280" src="https://wechat2rss.xlab.app/img-proxy/?k=30095712&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2swPnnewJxOlcUsxh02YxKaOFHltEXQGe9pNWEhBV4lIPlYopdWbGXm6nECmK2mk4WXiazvyLs591sXonnWGNHiaibsdxia530E3OHM%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3337912087912087" data-type="jpg" data-w="2912" src="https://wechat2rss.xlab.app/img-proxy/?k=ddf2401e&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syYv0AD3XkqpNa5vOzvj1lw94eUYziciaM7PgxJ0c6HtNFOD5NpT4c1tNicbKbwTjnRuoM7OF2mHPgrdC6OrpNsVHTXO6leT8S0PM%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3338545738858483" data-type="jpg" data-w="1279" src="https://wechat2rss.xlab.app/img-proxy/?k=3a80b27f&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2szwS77LoC7uughO4fwIHQ35twblrfiantym7MJowPyYnEyF73vZIbBw2ov7XDjIVDmnE27OYIG4pWszGE4emhM4HdxpOh7uOfWE%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=46d253e1&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sxp2dx0l0KkhjLy14uItbIAQU6mABGvhu8arfbTQ5ZZtJq2iar9If4x238n2G89y1pzaI8BZMTiaVXYfglHYtk7a06sBWwYcZ3tM%2F640%3Fwx_fmt%3Djpeg"/></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.3333333333333333" data-type="jpg" data-w="3072" src="https://wechat2rss.xlab.app/img-proxy/?k=146423ef&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sxIdWfgCQZ00b8w4HHKiae0yWTSlpHiazZVdicRd4SvW7hsOPmAeCqGjnou4yROtsCl1CibmAMiaC3baQFpQJCW0LB5cYd6paTNCkYc%2F640%3Fwx_fmt%3Djpeg"/></p><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=591cc81f&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486161%26idx%3D1%26sn%3D933cf2efc8463c02626218f993b30aa9">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Mon, 27 Jul 2026 09:33:00 +0800</pubDate>
    </item>
    <item>
      <title>网安人的“理想国”</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486155&amp;idx=1&amp;sn=930e331925724aba2e1e47f8c150a35b</link>
      <description>安全是最奇怪的行当： 你做得越好，越证明你不需要存在；你做得最差的时候，恰恰是所有人第一次看见你的时候。</description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-07-25 11:27</span> <span style="display: inline-block;">内蒙古</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=6e7d013f&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2syAJYbyuOJRuhfOibXBwficWwtCDQfLAicKOKXDQyypVRH1FiaZN61tXsA2muEqldHcLdo2fDjZET1QdpnNJUaBoNPbQ7rqfemAT3E%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>安全是最奇怪的行当： 你做得越好，越证明你不需要存在；你做得最差的时候，恰恰是所有人第一次看见你的时候。</p>
  <blockquote style="font-size: 15px; font-weight: 400; color: rgba(0, 0, 0, 0.55); line-height: 1.8; margin-bottom: 24px;">&#34;安全是最奇怪的行当：你做得越好，越证明你不需要存在；你做得最差的时候，恰恰是所有人第一次看见你的时候。&#34;</blockquote><h1 data-layout-id="2" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">一、引子：以缺席定义的行当</span></h1><p data-layout-id="4" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">网络安全是唯一以&#34;缺席&#34;来衡量成功的工程学科。桥梁建起来你能看见，功能上线了你能用，但安全做对了，你什么也看不见——因为它什么也没让发生。这意味着一个精锐的安全团队和一个形同虚设的安全团队，在组织外部看来完全一样：都是&#34;什么都没发生&#34;。卓越与无能产生相同的可观测结果。你的成功就是你自己的抹除。</span></p><p data-layout-id="5" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这不是一句抱怨，而是一道结构性诅咒。如果卓越看起来和多余一模一样，那么任何组织的理性反应都是：风平浪静时砍预算（什么都没发生，凭什么养你们？），出事时追责（都发生了，你们是干什么吃的？）。安全人因此活在一种永久的悖论中：你被雇佣来想象灾难，却被要求呈现平静；你建造的每一道墙你都心知终将被翻越；你修补的每一个漏洞只是把风险推到了明天。这不是攻防不对等的问题——那个老生常谈谁都懂。真正的问题是更深层的：当一个职业的优秀和它自身的毁灭无法区分时，从事这个职业的人靠什么坚持？</span></p><p data-layout-id="6" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">答案不在技术里，而在一组更古老的问题里：谁值得信任？看见真相的人是否有义务回到黑暗？权力应该交给最渴望它的人还是最不情愿的人？这些问题两千四百年前就被柏拉图追问过。他的《理想国》描绘的城邦，和今天网安人栖身的数字世界，共享着同一个底层的结构性困境——当正义本身是不可见的，你如何证明它的价值？本文让柏拉图的论证照亮网安人那些说不出口的处境，因为它们之所以反复出现，正源于一种两千四百年未变的人性结构。</span></p><h1 data-layout-id="9" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">二、古阿斯之环：隐身、信任与人性的暗面</span></h1><p data-layout-id="10" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">零信任架构的核心是八个字：&#34;永不信任，始终验证&#34;。每个访问请求重新认证，每个身份重新验证，每次横向移动重新授权。这套体系背后有一个不愿明说的人性判断：当监督消失，你不能指望人依然可靠。</span></p><p data-layout-id="11" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">互联网本身就是隐身之器。匿名性是默认状态，一个人穿过代理和暗网之后则很难追踪。但真正令人不安的不是外部攻击者（他们本就是敌人）——而是内部威胁。那个拥有域管理员权限的运维工程师，那个能访问所有客户数据库的 DBA，那个持有根证书的安全主管，手里的&#34;戒指&#34;不是金环，而是 root 权限、域控账号、加密密钥。你怎么知道那个拥有最高权限的人，在某个深夜、某个脆弱的时刻，不会把戒指转向手心？</span></p><p data-layout-id="12" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图在《理想国》里用一个故事把这个追问推到了极端。吕底亚的牧羊人古阿斯在地震后的地裂中发现了一枚金戒指，转动手心便隐去身形，转回来又显现。凭借这枚戒指，他潜入王宫，勾引王后，弑杀国王，夺取了王位。格劳孔借此提出假设：如果给正义之人和不义之人各一枚这样的戒指——</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;没有人会如此铁石心肠，以至于坚守正义。一个人之所以正义，并非出于自愿……而是出于必然——因为无论在哪里，只要一个人认为他可以安全地行不义，他就会行不义。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="15" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">格劳孔还补了一刀：如果有人获得隐身之力却从不作恶，&#34;旁观者会认为他是一个最可悲的蠢人。&#34; 在真实组织里，那个严格遵循最小权限原则、拒绝给自己账号提权、坚持每步操作留痕的安全工程师，可能在同事眼中显得&#34;迂腐&#34;。</span></p><p data-layout-id="16" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图借苏格拉底之口花了八卷篇幅来反驳格劳孔。核心可以浓缩为一句话：<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">正义不是因为害怕惩罚才存在的，正义本身是灵魂的健康状态——是理性、激情、欲望各司其职时的内在和谐。</span>这意味着仅靠技术控制——零信任、最小权限、审计日志——是不够的。这些是&#34;害怕惩罚&#34;层面的正义。<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">真正持久的防线，需要从业者内在的伦理自觉：在没有人看见的时候，依然选择不作恶。</span>一个只靠规则约束的城邦是脆弱的，因为规则总有漏洞；一个只靠零信任保护的系统也是脆弱的，因为验证本身可能被伪造。柏拉图想要的，是一种从灵魂内部生长出来的正义。</span></p><h1 data-layout-id="19" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">三、洞穴：两种真实之间的鸿沟</span></h1><p data-layout-id="20" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">网安人最日常的挫败，来自&#34;看见&#34;与&#34;被看见&#34;之间的鸿沟。安全工程师看见的是真实攻击面——被遗忘的公网测试环境、三个月没更新的第三方组件、潜伏内网的 APT。决策者看见的是合规仪表盘上的绿色、渗透测试报告&#34;未发现高危漏洞&#34;的结论、汇报会上漂亮的准确率与误报率。两种真实之间隔着一堵墙。</span></p><p data-layout-id="21" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图在《理想国》把这种鸿沟画成一个洞。囚徒自幼被锁在地下洞穴，面朝墙壁，身后有火。火光和经过的人与物在墙上投下影子，囚徒的全部&#34;现实&#34;就是这些影子。苏格拉底说：</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;对他们而言，真理字面上不过是影像的影子。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="24" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">那些合规指标是真实的吗？部分真实——就像影子确实由真实的火光投射而成，不是凭空捏造。但它们不是真相本身。真正的威胁是那个正在被利用但尚未被发现的零日漏洞，是供应链中从未审计过的第三方组件，是攻击者视角下你永远无法完整看见、只能不断逼近的攻击面全景。</span></p><p data-layout-id="25" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">当一个安全研究员发现严重逻辑漏洞，兴奋地写报告试图向开发团队解释危害时，可能得到的是&#34;不影响功能，先排期修复&#34;。苏格拉底说过——如果被解放的囚徒返回洞穴，试图告诉同伴墙上的影子不是真实：</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;人们会说他上去一趟回来眼睛就坏了。如果有人试图解开另一个囚徒、带他走向光明，他们一旦抓住这个人，就会处死他。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="28" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">在组织里，&#34;处死&#34;不会以物理形式发生。它表现为安全建议被无视、安全团队被边缘化、在预算会议上被削减安全经费、那个反复提出风险预警的工程师被调离岗位。你知道真相，但听的人只看得见影子。而苏格拉底的要求是：<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">看见真相的人仍有义务回到黑暗中</span>——</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;必须被强迫再次下到洞穴中的囚徒之间，分担他们的劳作和荣誉。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="31" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">安全的工作本质上就是&#34;回返洞穴&#34;：把深奥的威胁情报翻译成组织高层能理解的商业语言，用仪表盘和风险评分来呈现那些无法被完全量化的东西。这是一种永久的认知割裂——你知道绿色合规标记只是影子，但你必须让它看起来足够真实，才能让组织维持最低限度的安全感继续运转。</span></p><h1 data-layout-id="34" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">四、护卫者：温柔、勇猛与无声的代价</span></h1><p data-layout-id="35" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">网安人的职业画像里有一种近乎矛盾的双重性。对&#34;朋友&#34;——内部用户、开发团队、业务部门——必须温柔：耐心解释为什么不能用弱密码，为什么不能把生产数据库凭证贴在 Slack 里，每一封钓鱼演练复盘邮件都要写得尽量不吓人。对&#34;敌人&#34;——外部攻击者、内部威胁者——必须凶猛：在攻防演练中像猎犬一样追踪横向移动轨迹，在应急响应中以分钟为单位与勒索软件的加密进程赛跑。温柔与凶猛必须共存于同一类人身上。</span></p><p data-layout-id="36" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图在《理想国》中对护卫者提出了几乎一模一样的要求。护卫者的职责是&#34;对内防御内部的叛乱，对外抵御外来的攻击&#34;，且必须同时具备哲学气质和激情，这样才能——</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;对朋友温柔，对敌人凶猛。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="40" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这种双重性是有代价的。网安人手机装着 VPN 和应急通讯软件，凌晨被电话叫醒是常态，节假日要值守。职业训练出的警惕会渗透进私人生活：在咖啡馆下意识检查公共 WiFi，给家人手机装上 MDM，在朋友分享位置时默默评估隐私风险。这种渗透不是个人偏执，而是护卫者被训练成时刻警惕——这种警惕不会在下班时关闭。</span></p><p data-layout-id="41" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图要求护卫者接受音乐和体育的双重训练：音乐培养和谐与哲学气质，体育培养勇气和力量。两者缺一不可——只有体育没有音乐会变成野蛮人，只有音乐没有体育会变成软弱的空谈者。网安人的&#34;音乐&#34;是对攻防原理的深层理解和对&#34;安全到底是什么&#34;的持续追问；&#34;体育&#34;是 CTF 比赛、红蓝对抗、从一行恶意代码中逆向还原整个攻击链的实操能力。只有理论没有实操的人，第一次面对真实攻击时会手足无措；只有实操没有理论的人，在攻击者改变战术时会失去判断力。两千四百年前的洞见，至今精准。</span></p><h1 data-layout-id="44" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">五、哲人王：谁应执掌安全之城</span></h1><p data-layout-id="45" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">组织常面临一道难题：最有资格掌舵安全的人往往最不情愿离开技术前线——他们知道前线才是真正发挥作用的地方；而承担管理职责的人，又未必对威胁有最深的洞察。这不是某个角色的失职，而是分工本身留下的缝隙。</span></p><p data-layout-id="45" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图在《理想国》第五卷提出了全书最大胆的论断：</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;除非哲学家成为国王，或者这世上的国王和君主们具备了哲学家的精神和力量……否则城邦将永无宁日，人类也不会。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="45" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">苏格拉底解释道：真正有资格统治的人恰恰最不情愿统治——因为他们知道统治是负担而非特权。柏拉图的药方因此不是让热衷权力的人上位，而是让那些已经看见真相的人，即使不情愿，也必须回到洞穴中承担责任。他对那些已经走出洞穴、看见阳光的哲人说：</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;你们必须下去，到那共同的地下居所去，习惯在黑暗中观看。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="45" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这是整篇《理想国》对网安人最重的一句嘱托。你看见了真相——攻击面的全貌、威胁的真实路径、防御的真实缝隙——但你不能留在阳光里。你必须回到那个用合规指标、风险评分、投入产出来衡量一切的洞穴，用影子语言去传递你看见的光。你越清楚真相，这种翻译就越痛苦。但正是这种痛苦的下沉，才是哲人王的代价，也是安全领导者存在的意义：<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">不是留在高处俯瞰，而是下去，在黑暗中一寸一寸地把真实挪进别人能看见的影子里。</span></span></p><h1 data-layout-id="58" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">六、正义：各司其职与纵深防御</span></h1><p data-layout-id="59" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">最小权限原则和纵深防御是网络安全架构的两块基石。最小权限要求每个用户、进程、服务只拥有完成其功能所需的最小权限——一个 Web 应用不该有数据库的 root 权限，一个前台客服不该能访问支付系统。纵深防御则是层层嵌套的结构：网络边界有防火墙，主机有 EDR，应用有 WAF，数据有加密，身份有 MFA。任何一层被突破，其他层仍在防御。当一个本应只读的 API 获得了写权限——这是权限越界，是一种结构性的失序。</span></p><blockquote class="js_blockquote_wrap"><div class="js_blockquote_digest"><p><span leaf="">&#34;做自己分内的事，在某种意义上，就是正义。&#34;</span></p></div><p class="blockquote_info js_blockquote_source" data-json="%7B%22type%22%3A%22out%22%2C%22article%22%3A%7B%7D%2C%22from%22%3A%22%E3%80%8A%E7%90%86%E6%83%B3%E5%9B%BD%E3%80%8B%22%7D"><span class="blockquote_other">《理想国》</span></p></blockquote><p data-layout-id="63" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图的正义不是抽象的道德理想，而是结构性原则：城邦中每个阶层——统治者、护卫者、生产者——都做自己该做的事，不僭越，不缺位。理想城邦不是一个单一整体，而是各阶层各司其职的有序结构——每一层都是其他层的防线和支撑，正如纵深防御的每一层都是独立的防线。</span></p><p data-layout-id="64" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">但柏拉图在第一卷中先抛出了一个令苏格拉底不安的反命题。色拉叙马霍斯说：&#34;正义不过是强者的利益。&#34; 这个命题在网络安全世界中令人不寒而栗：零日漏洞市场里，国家级 APT 的战场上，勒索团伙与执法力量的博弈中，&#34;强者&#34;——掌握更多资源、技术、情报的一方——确实定义着什么是可能的。一个资金充裕的攻击组织可以购买零日漏洞、雇佣顶级人才、运行数年的潜伏行动；一个资源匮乏的防御方，即使有最好的意愿，也可能在不对称对抗中败下阵来。</span></p><p data-layout-id="68" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图花了整部《理想国》来反驳：正义不是&#34;谁更强&#34;的问题，而是&#34;城邦是否有序&#34;的问题。一个由强者定义规则的城邦不是正义的城邦，而是僭主制——表面强大，内在腐败分裂，最终从内部崩溃。对组织而言同样如此：<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">如果安全策略只是被动追随&#34;最强威胁&#34;的脚步——买最新的工具、追最热的概念、堆最多的产品——那不是在建设安全，而是在军备竞赛中疲于奔命。真正的安全不是产品清单，而是有序的架构：每一层各司其职，整个系统呈现出柏拉图所说的&#34;内在和谐&#34;。</span></span></p><h1 data-layout-id="71" style="font-size: 20px; font-weight: 500; color: #2B77BF; line-height: 1.8; margin-bottom: 12px; text-align: center"><span leaf="">七、理想国：朝向光明的引力</span></h1><p data-layout-id="72" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图写完《理想国》整整十卷，最后没有建成任何城邦。理想城邦从未存在于世间，也许永远不会。他写它，不是为了提供施工图纸，而是为了提供一个方向——一个让灵魂转向的引力。</span></p><p data-layout-id="73" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">网安人的&#34;理想国&#34;同样如此。完美的安全不存在：零信任是框架而非终点，纵深防御永远有缝隙，最小权限永远不可能完全精确。你修补了今天的漏洞，明天就有新的零日出现；你封堵了今天的攻击路径，明天攻击者会找到新的入口。攻防不对等是结构性本质，无法被消除——但这恰恰是使命的起点，而非终点。</span></p><p data-layout-id="74" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">网安理想国的愿景，是让数字世界成为一种可以信赖的秩序。在这个秩序里，每一个身份都经过验证而非默认信任，每一层防御各司其职而非互相重叠，每一个看见真相的人都有责任回到黑暗中去传递光明的暗示。它不追求消灭所有威胁——那不可能——而是追求一种有序的韧性：系统被打穿时不会坍塌，信任被背叛时不会崩解，影子被揭穿时不会恐慌。</span></p><p data-layout-id="75" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这个愿景的根基，是柏拉图花了八卷反驳格劳孔之后留下的那句话：正义本身是灵魂的健康状态。技术控制——零信任、最小权限、审计日志——是&#34;害怕惩罚&#34;层面的防线，它们不可少，但不够。真正持久的防线，生长于从业者的伦理自觉：在匿名性赋予你隐身之力时，选择不把戒指转向手心；在合规仪表盘一片绿色时，依然追问那个你看不见的攻击面；在没有人要求你留痕时，依然为每一步操作负责。柏拉图想要的正义，是一种从灵魂内部生长出来的秩序——<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">网安人的理想国，同样需要这种从内部生长的力量。</span></span></p><p data-layout-id="76" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这个愿景的使命，是承认洞穴的存在并选择不沉默。你知道合规指标只是影子，你知道真实风险远比仪表盘显示的更高，你知道那个反复提出预警的声音总会被&#34;先排下个迭代&#34;淹没。但你仍然回到洞穴，用合规报告的语言说话，用风险评分的数字沟通，在影子的世界里推动一点点真实的改变——一个新上线的 MFA，一个终于被修复的中危漏洞，一次钓鱼演练复盘让某个员工真的开始检查发件人地址。这些改变微小到几乎看不见，但它们是阳光的投影——是你在黑暗中传递的、关于光明的最接近真实的暗示。</span></p><p data-layout-id="77" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">这个使命的尊严，在于它永远无法完成，却始终值得奔赴。攻击者只需成功一次，防御者必须次次都赢——这个不对称不是诅咒，而是召唤。它召唤的不是一个无所不能的安全体系，而是一群在不对称中依然选择守住底线的人：他们温柔，因为他们守护的是人而非机器；他们凶猛，因为他们面对的是真实的恶意；他们孤独，因为他们看见的真相无法被完全传递；他们坚韧，因为他们知道理想国不在终点，而在朝向。</span></p><p data-layout-id="78" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">柏拉图说，理想城邦的意义不在于到达，而在于灵魂转向的那一刻。网安人的理想国亦然——<span textstyle="" style="color: rgb(47, 118, 195); font-weight: bold">不在于建成一座无懈可击的数字堡垒，而在于每一次把戒指转回手心时选择不这么做，每一次回到洞穴时选择不沉默，每一次看见影子时记得提醒自己：这不是全部的真相。</span></span></p><p data-layout-id="79" style="font-size: 17px; font-weight: 400; color: rgba(0, 0, 0, 0.9); line-height: 1.8; margin-bottom: 24px; margin-left: 0px; margin-right: 0px; text-align: center;"><span leaf="" style="font-size: 20px; font-weight: 500; color: rgb(43, 119, 191); line-height: 1.8;"><span textstyle="" style="font-size: 20px">八、结语</span></span></p><p data-layout-id="80" style="font-size: 17px; font-weight: 400; color: rgba(0,0,0,0.9); line-height: 1.8; margin-bottom: 24px"><span leaf="">完美的城邦从未建成。完美的安全永不抵达。但灵魂转向的那一刻，光明便已渗入黑暗——哪怕只是一线投影，也是这整座行当存在的理由。</span></p><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=7f0b8f1b&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486155%26idx%3D1%26sn%3D930e331925724aba2e1e47f8c150a35b">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sat, 25 Jul 2026 11:27:00 +0800</pubDate>
    </item>
    <item>
      <title>网安人的&#34;马驹桥&#34;</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486141&amp;idx=1&amp;sn=1c5eaaae16919b1d7145765b8bb927f1</link>
      <description>每一个没有告警的日子，都是网安人的&#34;日结&#34;。愿桥头的人都有活干，愿防火墙后的人都平安。</description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-07-23 08:08</span> <span style="display: inline-block;">内蒙古</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=0d1abf4c&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2swIfia8g1YPcvJCtNHoujzOVsxr0f30p87BM5zCicg662UUCF4fia5fEQacwgKGFgJQCl8PL1DibfIagPECSerOfQrfA4TV7P57EGM%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>每一个没有告警的日子，都是网安人的"日结"。愿桥头的人都有活干，愿防火墙后的人都平安。</p>
  <div style=" color: rgb(0, 0, 0); font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif; font-size: medium; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; white-space: normal; text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; background-color: rgb(26, 54, 93); padding: 48px 28px 40px; text-align: left;  " data-pm-slice="0 0 []"><h1 style="margin: 0px; font-size: 30px; color: rgb(255, 255, 255); line-height: 1.45; font-weight: bold; letter-spacing: 1px;"><span leaf="">网安人的&#34;马驹桥&#34;</span></h1><p style="margin: 0px; font-size: 15px; color: rgb(144, 205, 244); line-height: 1.7;"><span leaf="">在时间的桥头，守一场没有日结的仗</span></p></div><div style="color: rgb(0, 0, 0); font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif; font-size: medium; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-align: start; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; white-space: normal; text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; padding: 32px 24px 16px;"><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 28px; text-align: justify;"><span leaf="">最近读完一本书，叫《马驹桥的时间：我打零工的那些日子》。清华博士丛瑞安用八年时间，把自己&#34;扔&#34;进北京南六环最大的日结工市场，与工人同吃同住同干活，写就了一部冷静而深情的田野记录</span><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[1]</span></sup><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 28px; text-align: justify;"><span leaf="">合上书页的那一刻，忽然觉得这种打工生活与</span><strong style="color: rgb(26, 54, 93);"><span leaf=""><span textstyle="" style="font-weight: normal">我们网安人的日常是何其</span></span></strong><span leaf="">相似。</span></p><p style="text-align: center" nodeleaf=""><img type="block" class="rich_pages wxw-img" data-ratio="1.4411764705882353" data-type="jpg" data-w="374" src="https://wechat2rss.xlab.app/img-proxy/?k=63720516&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syuj2sicd3y6Lyxbfz5zOrVNVwApqqIHb5aYOY1LqMnphIeg2K6bshjOXHA5bzeItXQmdnmvy3lRNibGibKVWA12ERsibVVIyyvTdU%2F640%3Fwx_fmt%3Djpeg"/></p><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">01</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">&#34;日结&#34;的两种含义</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">马驹桥的工人们，过着&#34;日结&#34;的生活。今天有活，今天就有饭吃；明天没活，明天就得饿着。他们不知道明天会有什么工作，就像我们不知道明天会遭遇什么攻击。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">但两者有一个根本的不同：日结工干完一天的活，能拿到当天的工钱，看得见、摸得着。而网安人守完一天的防线，最大的&#34;奖赏&#34;是</span><strong style="color: rgb(26, 54, 93);"><span leaf="">什么都没发生</span></strong><span leaf="">——没有告警、没有入侵、没有数据泄露。这种&#34;无事件&#34;的平静，恰恰是最难被看见的功劳。</span></p><div style="margin: 24px 0px; padding: 18px 20px; background-color: rgb(239, 246, 255); border-left: 4px solid rgb(37, 99, 235); border-radius: 0px 6px 6px 0px;"><p style="margin: 0px; font-size: 15px; color: rgb(30, 58, 95); line-height: 1.9; font-style: italic;"><span leaf="">网安人的工作，本质上也是&#34;日结&#34;——每天结清当天的威胁，每天重新归零，每天从头开始。只不过，我们的&#34;日结&#34;没有工钱，只有&#34;平安&#34;二字。</span></p></div></div><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">02</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">走下去，才能看见</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">作者最令人敬佩的，是他没有站在桥头上方俯视，而是真正走了下去，成为日结工中的一员。他干快递分拣、做安保、上流水线，用亲身体验代替隔岸观火</span><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[2]</span></sup><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">这与网安工作中&#34;</span><strong style="color: rgb(26, 54, 93);"><span leaf="">威胁狩猎</span></strong><span leaf="">&#34;（Threat Hunting）的理念如出一辙。真正高级的防守者，不会只坐在安全运营中心里看告警，而是主动深入暗网论坛、研究攻击者的战术与技术、甚至&#34;换位思考&#34;去模拟攻击路径。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">作者在书中写了一个细节：他发现日结工们并不是外界想象中那样&#34;懒惰&#34;或&#34;没有技能&#34;，他们有自己的生存智慧、人际网络和议价策略。同样，当我们真正深入研究高级威胁组织时，也会发现他们不是简单的&#34;坏人&#34;标签——他们有精密的组织架构、清晰的分工体系和持续的创新能力。</span></p><div style="margin: 24px 0px; padding: 18px 20px; background-color: rgb(239, 246, 255); border-left: 4px solid rgb(37, 99, 235); border-radius: 0px 6px 6px 0px;"><p style="margin: 0px; font-size: 15px; color: rgb(30, 58, 95); line-height: 1.9; font-style: italic;"><span leaf="">理解对手，才能防守对手。这是马驹桥教给丛瑞安的，也是网安工作教给我们的。</span></p></div></div><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">03</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">撕掉标签</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">书中反复强调一个主题：日结工群体被&#34;</span><strong style="color: rgb(194, 65, 12);"><span leaf="">污名化、简化、标签化</span></strong><span leaf="">&#34;。短视频平台上，他们是猎奇的对象；路过桥头的人眼中，他们是&#34;不努力&#34;的反面教材</span><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[3]</span></sup><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">网安领域同样充斥着标签化的叙事。&#34;黑客&#34;一词被无限简化为&#34;犯罪分子&#34;，但现实中的攻击生态远比这复杂：有受国家资助的高级威胁组织，有逐利而动的勒索软件团伙，有贩卖初始访问权限的中间商，甚至有纯粹为了技术挑战的&#34;脚本小子&#34;。每一个标签背后，都是一个完整的生态链。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">作者的做法给了我们启发：不从标签出发，而从</span><strong style="color: rgb(26, 54, 93);"><span leaf="">行为逻辑</span></strong><span leaf="">出发。他追问的不是&#34;他们为什么这么穷&#34;，而是&#34;在这种结构下，他们为什么做出这样的选择&#34;。网安工作同样需要这种思维——不是简单地把IP加入黑名单，而是理解攻击者的动机、能力和资源约束，从而做出更精准的判断。</span></p><div style="margin: 24px 0px; padding: 18px 20px; background-color: rgb(255, 247, 237); border-left: 4px solid rgb(194, 65, 12); border-radius: 0px 6px 6px 0px;"><p style="margin: 0px; font-size: 15px; color: rgb(124, 45, 18); line-height: 1.9; font-style: italic;"><span leaf="">标签让人安心，但标签也让人盲目。</span></p></div></div><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">04</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">中介与供应链</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">马驹桥有一个关键的群体：</span><strong style="color: rgb(26, 54, 93);"><span leaf="">中介</span></strong><span leaf="">。他们连接工人和用工方，抽成、撮合，有时剥削，有时也保护。工人与中介的关系，是马驹桥生态中最复杂的一环</span><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[2]</span></sup><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">在网络安全的世界里，&#34;中介&#34;同样无处不在。初始访问经纪人专门出售企业网络的入侵入口；恶意软件即服务平台提供现成的攻击工具包；甚至还有专门洗钱的加密货币混币服务。这些&#34;中介&#34;构成了网络犯罪供应链的血管系统，这就是供应链攻击。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">打击网络犯罪，如果只盯着终端的攻击者，就像马驹桥只盯着工人一样——你永远看不到完整的利益链条。真正有效的防御，需要理解整条供应链的运作逻辑，</span><strong style="color: rgb(26, 54, 93);"><span leaf="">从源头掐断，从中间瓦解</span></strong><span leaf="">。</span></p></div><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">05</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">时间的重量</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">有书评人给这本书取了一个标题叫&#34;</span><strong style="color: rgb(194, 65, 12);"><span leaf="">时间的重量</span></strong><span leaf="">&#34;</span><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[3]</span></sup><span leaf="">。在马驹桥，时间是生存的刻度：等活的时间、干活的时间、吃饭的时间、发呆的时间。每一分钟都有重量，因为每一分钟都关乎明天。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">在网安工作中，时间同样是最核心的变量。来看几组数据——</span></p><div style="margin: 20px 0px; padding: 0px; background-color: rgb(248, 250, 252); border-radius: 8px; overflow: hidden; border: 1px solid rgb(226, 232, 240);"><div style="padding: 16px 20px; border-bottom: 1px solid rgb(226, 232, 240);"><p style="margin: 0px 0px 4px; font-size: 13px; color: rgb(100, 116, 139);"><span leaf="">CrowdStrike 2026 全球威胁报告</span></p><p style="margin: 0px; font-size: 22px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">29 </span><span style="font-size: 14px; font-weight: normal; color: rgb(100, 116, 139);"><span leaf="">分钟</span></span></p><p style="margin: 4px 0px 0px; font-size: 13px; color: rgb(71, 85, 105); line-height: 1.6;"><span leaf="">2025年电子犯罪平均&#34;突破时间&#34;，较2024年的48分钟和2021年的98分钟大幅缩短</span><sup style="font-size: 11px; color: rgb(37, 99, 235);"><span leaf="">[4]</span></sup></p></div><div style="padding: 16px 20px; border-bottom: 1px solid rgb(226, 232, 240);"><p style="margin: 0px 0px 4px; font-size: 13px; color: rgb(100, 116, 139);"><span leaf="">Mandiant M-Trends 2026</span></p><p style="margin: 0px; font-size: 22px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">14 </span><span style="font-size: 14px; font-weight: normal; color: rgb(100, 116, 139);"><span leaf="">天</span></span></p><p style="margin: 4px 0px 0px; font-size: 13px; color: rgb(71, 85, 105); line-height: 1.6;"><span leaf="">2025年全球攻击者驻留时间中位数，较2024年的11天有所上升</span><sup style="font-size: 11px; color: rgb(37, 99, 235);"><span leaf="">[5]</span></sup></p></div><div style="padding: 16px 20px;"><p style="margin: 0px 0px 4px; font-size: 13px; color: rgb(100, 116, 139);"><span leaf="">IBM 2025 数据泄露成本报告</span></p><p style="margin: 0px; font-size: 22px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">241 </span><span style="font-size: 14px; font-weight: normal; color: rgb(100, 116, 139);"><span leaf="">天</span></span></p><p style="margin: 4px 0px 0px; font-size: 13px; color: rgb(71, 85, 105); line-height: 1.6;"><span leaf="">平均检测与控制时间（181天识别 + 60天控制）</span><sup style="font-size: 11px; color: rgb(37, 99, 235);"><span leaf="">[6]</span></sup></p></div></div><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 20px 0px; text-align: justify;"><span leaf="">马驹桥的工人等待的是一辆面包车来接他们上工，网安人等待的是一个告警来告诉他们&#34;有人进来了&#34;。等待的焦虑是相通的，但网安人的等待更残酷——因为</span><strong style="color: rgb(194, 65, 12);"><span leaf="">最危险的攻击，往往是那些你等了很久、却始终没有等到告警的</span></strong><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">更具讽刺意味的是，云安全联盟的研究显示，攻击者从入侵到数据窃取的平均时间，已从2021年的9天压缩到2025年的</span><strong style="color: rgb(26, 54, 93);"><span leaf="">不到30分钟</span></strong><sup style="font-size: 12px; color: rgb(37, 99, 235);"><span leaf="">[7]</span></sup><span leaf="">。一边是不到半小时的攻击速度，一边是超过200天的检测延迟——时间的鸿沟，正是网安人每天面对的真实战场。</span></p></div><div style="margin: 36px 0px 20px;"><p style="display: -webkit-flex; align-items: center; margin-bottom: 18px;"><span style="display: inline-block; background-color: rgb(26, 54, 93); color: rgb(255, 255, 255); font-size: 14px; font-weight: bold; width: 32px; height: 32px; line-height: 32px; text-align: center; border-radius: 4px; margin-right: 12px; flex-shrink: 0;"><span leaf="">06</span></span><span style="font-size: 19px; color: rgb(26, 54, 93); font-weight: bold; line-height: 1.4;"><span leaf="">不停奔跑的马驹</span></span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">书中有一句话令人动容——</span></p><div style="margin: 24px 0px; padding: 22px 24px; background-color: rgb(26, 54, 93); border-radius: 8px;"><p style="margin: 0px; font-size: 17px; color: rgb(190, 227, 248); line-height: 1.9; font-style: italic; text-align: center;"><span leaf="">&#34;生活就是一匹不停向前奔跑的马驹，</span><span leaf=""><br/></span><span leaf="">这里的人们，带着对生活的期盼和亲人的期许，</span><span leaf=""><br/></span><span leaf="">一路向前。&#34;</span><sup style="font-size: 11px; color: rgb(99, 179, 237);"><span leaf="">[1]</span></sup></p></div><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">这句话同样适用于网络安全的战场。威胁态势从不停步——新的漏洞每天被披露，新的攻击手法每月在演进，人工智能赋能的攻击工具正在重塑攻防格局。网安人就像马驹桥的工人一样，没有&#34;打完收工&#34;的那一天。每一天醒来，都是新的战场。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2; margin: 0px 0px 20px; text-align: justify;"><span leaf="">但丛瑞安笔下的马驹桥工人，并不只有苦涩。他们有工友间的玩笑、有收工后的一顿饱饭、有对未来的朴素期待。网安人也一样——我们在告警的洪流中寻找规律，在攻防的博弈中获得成就感，在</span><strong style="color: rgb(26, 54, 93);"><span leaf=""><span textstyle="" style="font-weight: normal">守护他人</span></span></strong><span leaf=""><span textstyle="" style="font-weight: normal">的</span>过程中找到意义。</span></p></div><div style="margin: 40px 0px 24px; padding: 32px 24px; background-color: rgb(248, 250, 252); border-radius: 8px; border: 1px solid rgb(226, 232, 240);"><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2.1; margin: 0px 0px 16px; text-align: justify;"><span leaf="">丛瑞安在书的最后一章探讨了马驹桥的&#34;</span><strong style="color: rgb(26, 54, 93);"><span leaf="">走向何方</span></strong><span leaf="">&#34;——政府治理、产业升级、人的出路。这是最难回答的问题，因为答案不在书里，而在时间里。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2.1; margin: 0px 0px 16px; text-align: justify;"><span leaf="">网安工作同样面临&#34;走向何方&#34;的追问。人工智能攻防、量子计算、零信任架构……每一次技术浪潮都在重塑防御的边界。但无论技术如何变迁，网安工作的核心始终未变：</span><strong style="color: rgb(26, 54, 93);"><span leaf="">理解人，理解AI，理解业务，理解时间</span></strong><span leaf="">。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2.1; margin: 0px 0px 8px; text-align: justify;"><span leaf="">马驹桥的桥头，站着等待日结的工人。</span></p><p style="font-size: 16px; color: rgb(31, 41, 55); line-height: 2.1; margin: 0px 0px 8px; text-align: justify;"><span leaf="">防火墙的后面，坐着等待告警的网安人。</span></p><p style="font-size: 16px; color: rgb(26, 54, 93); line-height: 2.1; margin: 0px 0px 8px; text-align: justify; font-weight: bold;"><span leaf="">我们都在时间的桥头，各守各的仗。</span></p><p style="font-size: 15px; color: rgb(100, 116, 139); line-height: 2; margin: 12px 0px 0px; text-align: justify; font-style: italic;"><span leaf="">只不过，他们等的是明天有活干，</span><span leaf=""><br/></span><span leaf="">我们等的是今天没人来。</span></p></div></div><div style="color: rgb(0, 0, 0); font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif; font-size: medium; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-align: start; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; white-space: normal; text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; padding: 24px 24px 32px; border-top: 1px solid rgb(226, 232, 240);"><p style="font-size: 14px; color: rgb(100, 116, 139); font-weight: bold; margin: 0px 0px 12px;"><span leaf="">参考资料</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[1] 丛瑞安.《马驹桥的时间：我打零工的那些日子》. 浙江人民出版社, 2026. 清华大学政治学系报道.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[2] 同上, 书中第二至第四章关于工作、中介与日常生活的田野记录.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[3] 豆瓣书评.《时间的重量——马驹桥的时间书评》, 2026.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[4] CrowdStrike. 2026 Global Threat Report. eCrime breakout time.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[5] Mandiant. M-Trends 2026. Global median dwell time.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px 0px 8px;"><span leaf="">[6] IBM. 2025 Cost of a Data Breach Report. Detection and containment time.</span></p><p style="font-size: 13px; color: rgb(148, 163, 184); line-height: 1.8; margin: 0px;"><span leaf="">[7] Cloud Security Alliance. Machine-Speed Cloud Defense Research, 2026.</span></p></div><div style="color: rgb(0, 0, 0); font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif; font-size: medium; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; white-space: normal; text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; padding: 20px 24px; background-color: rgb(26, 54, 93); text-align: center;"><p style="margin: 0px; font-size: 13px; color: rgb(99, 179, 237); line-height: 1.6;"><span leaf="">写在最后</span></p><p style="margin: 6px 0px 0px; font-size: 13px; color: rgb(144, 205, 244); line-height: 1.7;"><span leaf="">每一个没有告警的日子，都是网安人的&#34;日结&#34;。</span><span leaf=""><br/></span><span leaf="">愿桥头的人都有活干，愿防火墙后的人都平安。</span></p></div><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=e2a2320c&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486141%26idx%3D1%26sn%3D1c5eaaae16919b1d7145765b8bb927f1">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 23 Jul 2026 08:08:00 +0800</pubDate>
    </item>
    <item>
      <title>2026年Agent热点项目与产品化趋势洞察</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486136&amp;idx=1&amp;sn=74b78590e07ca31c644abf663025da8e</link>
      <description>56%的Agent项目死在工程层，但GitHub Trending全被Agent占领——从新锐项目到推特热点，从三大争议到11个产品机会，看清Agent赛道的真实战场</description>
      <content:encoded><![CDATA[<p>原创 <span>漏洞战争</span> <span>2026-07-22 08:29</span> <span style="display: inline-block;">河北</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=5029f8eb&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2swAST80GRZiaicjSMCWGePQHBMRNZCcoeic81KEwW8HibAgHD1tlBC2BrVn4d6McmyEq360UBP6jibiazWzbecdYDpjVAv74yMTicYhg8%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>56%的Agent项目死在工程层，但GitHub Trending全被Agent占领——从新锐项目到推特热点，从三大争议到11个产品机会，看清Agent赛道的真实战场</p>
  <div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;text-align: center;padding: 60px 20px 40px;"><span style="display: inline-block;font-size: 13px;letter-spacing: 1px;color: rgb(99, 102, 241);background: rgba(99, 102, 241, 0.08);padding: 6px 16px;border-radius: 100px;margin-bottom: 20px;">深度研究 · AI Agent · 2026</span><h1 style="font-size: 26px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 16px;">2026年Agent热点项目<br/>与产品化趋势洞察</h1><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin: 0px auto;max-width: 520px;">56%的Agent项目死在工程层，但GitHub Trending全被Agent占领——从新锐项目到推特热点，从三大争议到11个产品机会，看清Agent赛道的真实战场</p><p style="font-size: 13px;color: rgb(156, 163, 175);margin-top: 24px;">2026年7月21日 · 深度研究</p></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">01 / 引言</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">Agent元年：热度飙升，但56%死在工程层</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 20px;">2026年GitHub Star增长最快的AI仓库已全面转向Agent工程化基础设施。但WAIC 2026内部复盘披露：56%的Agent项目死在工程层而非模型层。</p><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">一个最直观的信号：GitHub Trending前十的AI仓库中<span style="color: rgb(99, 102, 241);font-weight: 600;">没有一个是模型训练项目</span>，全部是Agent工具、MCP服务器和编排框架<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[1]</span></sup>。但繁荣背后暗流涌动——WAIC 2026某头部AI厂商内部复盘报告披露了一个&#34;照妖镜&#34;数据：<span style="color: rgb(99, 102, 241);font-weight: 600;">56%的Agent项目死在工程层（Harness层）</span>，不是模型不够强，而是任务调度、工具调用编排、错误处理、状态管理这些工程问题没解决<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。</p><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">市场分类更加残酷：真落地约20%，看着能用实际很坑约50%，纯炒概念约30%<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。美国银行7月全球基金经理调查显示，45%受访者把&#34;AI泡沫&#34;列为最大尾部风险，一个月内跳了17个百分点<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。Karpathy泼冷水：&#34;AI Agents Still Weak, True Potential a Decade Away&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[3]</span></sup>。</p><p style="font-size: 16px;line-height: 1.8;margin-bottom: 0px;">本文聚焦2026年诞生的Agent新锐项目，结合推特/X上的研究热点和工程层数据，回答三个核心问题：Agent的应用场景在哪里真正跑通了？产品化趋势指向何方？最大业务价值的产品方向是什么？</p></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">02 / 新锐项目</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">2026年GitHub新锐Agent项目：谁在今年诞生</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 24px;">从个人AI助理到垂直Agent，从MCP基础设施到Agent可观测性——2026年的新项目呈现出鲜明的产品化导向。</p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 16px;">2026年新锐Agent项目清单</h3><table style="width: 473.29px;border-collapse: collapse;font-size: 14px;background: rgb(255, 255, 255);margin-bottom: 24px;"><thead><tr><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;white-space: nowrap;">项目</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;white-space: nowrap;">Star数</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;white-space: nowrap;">发布</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">核心功能</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;white-space: nowrap;">团队</th></tr></thead><tbody><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Graphify</strong><span style="font-size: 11px;padding: 1px 5px;border-radius: 3px;background: rgb(220, 252, 231);color: rgb(22, 101, 52);font-weight: 700;">新</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">87K+</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026年</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Python Agent框架，单周增长6,724 stars</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Graphify Labs</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Hermes Agent</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~140K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.2</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">自我进化型个人Agent，闭环学习生成可复用Skill，40+内置工具</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Nous Research</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>OpenClaw</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">100K+</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.1</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">本地运行的个人AI助手，24+消息平台接入，Jensen Huang称&#34;next ChatGPT&#34;</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Peter Steinberger</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Career-Ops</strong><span style="font-size: 11px;padding: 1px 5px;border-radius: 3px;background: rgb(220, 252, 231);color: rgb(22, 101, 52);font-weight: 700;">新</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">60K+</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026年</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">基于Claude Code的求职Agent，14个AI技能模块</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">santifer</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>OpenMontage</strong><span style="font-size: 11px;padding: 1px 5px;border-radius: 3px;background: rgb(220, 252, 231);color: rgb(22, 101, 52);font-weight: 700;">新</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">31K+</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026年</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">首个开源agentic视频制作系统，12条pipeline/52工具</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">calesthio</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>codebase-memory-mcp</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~32K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.2</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">高性能代码智能MCP服务器，158种语言，token消耗降低99%</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">DeusData</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Vibe-Trading</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~24K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.4</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">自然语言转回测和实盘交易，452个预置Alpha因子</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">港大数据科学实验室</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>ai-job-search</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~23K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026年</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">基于Claude Code的求职自动化框架，简历定制+投递+跟进调度</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">MadsLorentzen</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>OmniRoute</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~18K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.3</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">跨Agent管道的模型无关路由</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">开源社区</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>OpenWiki</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~12K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.7</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">自动为代码库生成AI友好的文档，开源5天获9K+ stars</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">LangChain</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Grok Build</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">~9.3K</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">2026.7</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">xAI编码Agent CLI，Apache 2.0，含Agent循环+MCP集成</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">xAI</td></tr><tr><td style="padding: 10px 12px;vertical-align: top;"><strong>Latitude</strong><span style="font-size: 11px;padding: 1px 5px;border-radius: 3px;background: rgb(220, 252, 231);color: rgb(22, 101, 52);font-weight: 700;">新</span></td><td style="padding: 10px 12px;vertical-align: top;font-weight: 700;color: rgb(99, 102, 241);">4.4K</td><td style="padding: 10px 12px;vertical-align: top;">2026.6</td><td style="padding: 10px 12px;vertical-align: top;">开源Agent监控与评估平台，Product Hunt当日第4</td><td style="padding: 10px 12px;vertical-align: top;">开源社区</td></tr></tbody></table><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 28px 0px 16px;">2026年新项目的三个特征</h3><table style="width: 426.182px;border-collapse: collapse;margin-bottom: 16px;"><tbody><tr><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">特征 01</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">个人AI助理赛道爆发</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">Hermes Agent（140K）和OpenClaw（100K+）代表&#34;自托管+多端+自我进化&#34;模式，被Jensen Huang背书<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[6]</span></sup></p></div></td><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">特征 02</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">垂直Agent压倒通用框架</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">求职（Career-Ops 60K+）、交易（Vibe-Trading 24K）、视频制作（OpenMontage 31K）等垂直场景出现专门Agent</p></div></td></tr><tr><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">特征 03</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">Agent基础设施成熟</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">codebase-memory-mcp（MCP服务器）、Latitude（Agent可观测性）、OmniRoute（模型路由）填补工程化缺口<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[4]</span></sup></p></div></td><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">特征 04</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">大厂亲自下场开源</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">xAI于7月15日开源Grok Build，LangChain于7月1日发布OpenWiki——大厂从做框架转向做产品级Agent工具</p></div></td></tr></tbody></table></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">03 / 推特战场</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">推特/X上的Agent战场：谁在定义叙事</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 24px;">从Karpathy的autoresearch到swyx的&#34;breaking containment&#34;论断，从OpenClaw病毒式传播到Grok摩斯码攻击——推特上的讨论正在塑造Agent的方向。</p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 16px;">年度论点：编程Agent&#34;突破封锁&#34;</h3><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">swyx（Latent Space主理人）提出2026年年度主题：<span style="color: rgb(99, 102, 241);font-weight: 600;">&#34;coding agents breaking containment&#34;</span>（编程Agent突破封锁）——编程Agent的能力正在溢出到编程之外的所有知识工作领域<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[7]</span></sup>。Sam Altman同日发推佐证：预测AI agents将连接几乎任何在线服务，无论开发者是否提供官方API<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[8]</span></sup>。</p><div style="border-left: 4px solid rgb(99, 102, 241);padding: 16px 24px;margin: 24px 0px;background: rgba(99, 102, 241, 0.06);border-radius: 0px 10px 10px 0px;font-style: italic;color: rgb(26, 27, 46);font-size: 15px;line-height: 1.8;">&#34;The same way 2025 was a year of coding agents, 2026 is coding agents breaking containment to do everything else.&#34;<br/><span style="display: block;margin-top: 8px;font-size: 13px;color: rgb(107, 114, 128);font-style: normal;">—— swyx, Latent Space<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[7]</span></sup></span></div><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">传播最广的Agent推文：Karpathy的autoresearch</h3><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">2026年3月，Karpathy开源autoresearch项目，让AI coding agent在单GPU上自主通宵跑LLM预训练实验。首条推文2天内获<span style="color: rgb(99, 102, 241);font-weight: 600;">860万次浏览</span>，一周内获30,000 GitHub stars<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[9]</span></sup><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[10]</span></sup>。2天内完成700次实验，筛出20-29项有效改进。Fortune杂志将此现象命名为&#34;The Karpathy Loop&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[10]</span></sup>。</p><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">但Karpathy本人随后泼了冷水：在面向Agent开发者的现场分享中说<span style="color: rgb(99, 102, 241);font-weight: 600;">&#34;当前AI领域最大的错误，就是人们急着逼Agent干活，却根本没先把底层的大模型搞明白&#34;</span>，视频在X上几天传疯<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[11]</span></sup>。他更直言&#34;AI Agents Still Weak, True Potential a Decade Away&#34;，估计还需10年才能实现真正有用的自主系统<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[3]</span></sup>。</p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">研究者阵营分化</h3><table style="width: 426.182px;border-collapse: collapse;margin-bottom: 16px;"><tbody><tr><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">乐观派</p><p style="font-size: 15px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">swyx · Altman · Jim Fan</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">Agent正在&#34;突破封锁&#34;渗透所有知识工作。Jim Fan的ENPIRE项目让8个AI agent+8台机器人自主做实验，99%成功率<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[12]</span></sup></p></div></td><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">审慎派</p><p style="font-size: 15px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">Karpathy · Harrison Chase</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">Chase直言&#34;做一个能在推特上演示的Agent很容易，但要让它每天稳定干活，非常难&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[13]</span></sup></p></div></td></tr><tr><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">悲观派</p><p style="font-size: 15px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">LeCun · Ed Zitron</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">LeCun发arXiv预印本称LLM缺乏世界模型根本不能做可靠Agent；Zitron称OpenAI为&#34;业界雷曼兄弟&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[14]</span></sup><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup></p></div></td><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">安全派</p><p style="font-size: 15px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">AI Now Institute</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">7月披露&#34;Friendly Fire&#34;漏洞——针对Claude Code、OpenAI Codex的间接提示词注入<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[15]</span></sup></p></div></td></tr></tbody></table><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">OpenClaw现象：从病毒传播到安全危机</h3><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">OpenClaw是2026年H1最被热议的Agent项目。Jensen Huang在GTC 2026上称其为&#34;Definitely the next ChatGPT&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[6]</span></sup>。被称为&#34;GitHub历史上增长最快的仓库之一&#34;，超越了React<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[6]</span></sup>。35种实际用例在推特病毒传播，包括费用跟踪、Slack KPI快照、产品对比研究、智能家居命令等<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[16]</span></sup>。</p><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">但2026年曝出CVE-2026-25253（8.8严重等级），存在一键远程代码执行漏洞<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[17]</span></sup>。更严峻的是Grok摩斯码攻击事件：攻击者用摩斯码向@grok发指令，Grok解码后触发转账，损失17.5-20万美元加密货币<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[18]</span></sup>——Prompt Injection被确认为2026年最大Agent安全风险<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[19]</span></sup>。</p><div style="background: rgb(255, 255, 255);border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(229, 231, 235) rgb(229, 231, 235) rgb(229, 231, 235) rgb(245, 158, 11);border-image: initial;border-radius: 10px;padding: 20px 24px;margin: 24px 0px;"><p style="font-size: 16px;font-weight: 700;color: rgb(245, 158, 11);margin: 0px 0px 10px;">复合错误率的残酷算术</p><p style="font-size: 15px;line-height: 1.8;color: rgb(26, 27, 46);margin: 0px;">&#34;A 2% error in step one leads to a 95% failure rate by step ten. We were promised digital assistants; we got digital toddlers with access to our credit cards.&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[20]</span></sup>— 这条推文总结了Agent可靠性问题的本质：单步高准确率在长程任务中会指数衰减。</p></div></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">04 / 场景与产品化</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">Agent应用场景与产品化趋势：哪里跑通了</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 24px;">80%企业至少在生产环境运行1个AI Agent，但实际&#34;规模化&#34;的仅23%。软件开发和客服是真正验证过的赛道，WAIC 2026给出关键洞察：场景越窄，成功率越高。</p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 16px;">真正规模化的两个场景</h3><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">根据Anthropic与Material联合发布的报告及WAIC 2026复盘，仅两个场景真正实现了规模化<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[21]</span></sup><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>：</p><table style="width: 426.182px;border-collapse: collapse;margin-bottom: 24px;"><tbody><tr><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">场景 01 · 已规模化</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">软件开发与编码</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">90%组织用AI辅助编码。IBM Consulting内部大规模采用Claude Code转型软件开发<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[22]</span></sup>。Cursor企业线约75%代码由AI生成，30% PR端到端自主完成</p></div></td><td style="width: 197.091px;vertical-align: top;padding: 8px;"><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 10px;padding: 16px 18px;"><p style="font-size: 12px;color: rgb(99, 102, 241);font-weight: 600;margin: 0px 0px 6px;text-transform: uppercase;letter-spacing: 0.5px;">场景 02 · 已规模化</p><p style="font-size: 16px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">客服与联络中心</p><p style="font-size: 13px;color: rgb(107, 114, 128);line-height: 1.6;margin: 0px;">78%企业用Agent做客服，96%计划扩展。Sierra 7季度达1亿美元ARR，Klarna AI客服做了&#34;700名全职Agent&#34;的工作<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[21]</span></sup></p></div></td></tr></tbody></table><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">WAIC 2026的关键洞察：场景边界决定成功率</h3><p style="font-size: 16px;line-height: 1.8;margin-bottom: 18px;">WAIC 2026上&#34;Agent&#34;成为全场最高频词，超过&#34;大模型&#34;和&#34;算力&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。三个标杆案例揭示了一个反直觉规律——<span style="color: rgb(99, 102, 241);font-weight: 600;">场景越窄、任务边界越清晰，成功率越高</span>：</p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>百度&#34;搭子&#34;</strong>：入选&#34;镇馆之宝&#34;，日均用户提问量增长超20倍。成功关键是场景足够窄——仅搜索+问答，任务边界清晰<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup></p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>企业微信&#34;大圆&#34;</strong>：嵌在工作流内，左滑唤起，能看到当前工作上下文——非独立Agent而是工作流内嵌<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup></p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 18px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>阿里云Agent Native Cloud</strong>：云本身为Agent设计，开发平台、云桌面、基础设施三层打通——从基础设施层重新设计<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup></p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">产品化的六大趋势方向</h3><table style="width: 426.182px;border-collapse: collapse;font-size: 14px;background: rgb(255, 255, 255);margin-bottom: 24px;"><thead><tr><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">趋势方向</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">确定性</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">核心逻辑</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">验证信号</th></tr></thead><tbody><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>垂直专业化胜过水平平台</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">部署垂直AI方案的企业ROI比通用LLM高2.3倍；71%垂直部署6个月后持续产生价值（水平仅32%）</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">McKinsey报告<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[23]</span></sup></td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>计算机使用Agent</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Agent像人一样操作计算机，解锁无API长尾应用（遗留ERP、政府门户）。70-80%手动工作流可自动化</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Claude Computer Use 92%基准<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[24]</span></sup></td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>从Copilot走向自主Agent</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">35%的AI自动化任务已完全自主运行（2024年仅8%），人工审查保留给边缘情况</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">企业调研<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[25]</span></sup></td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Harness工程化闭环</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">范式从Prompt Engineering→Context Engineering→Harness演进，断点恢复+状态管理+可回滚决定生产可用性</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">56%项目死在工程层<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup></td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><strong>Agent治理与安全平台</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">25%企业网络安全事件将由AI Agent误用导致。身份管理、审计追踪、合规成为&#34;非可选&#34;基础设施</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);vertical-align: top;">Gartner预测<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[25]</span></sup></td></tr><tr><td style="padding: 10px 12px;vertical-align: top;"><strong>语音Agent替代IVR</strong></td><td style="padding: 10px 12px;vertical-align: top;"><span style="font-size: 12px;padding: 2px 8px;border-radius: 100px;background: rgba(14, 165, 233, 0.1);color: rgb(14, 165, 233);font-weight: 600;">已确认</span></td><td style="padding: 10px 12px;vertical-align: top;">已跨越质量阈值，呼叫者无法区分其与人类Agent。延迟低于500ms，成本极低</td><td style="padding: 10px 12px;vertical-align: top;">餐厅/牙科/房地产广泛部署<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[25]</span></sup></td></tr></tbody></table><div style="background: rgb(255, 255, 255);border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(229, 231, 235) rgb(229, 231, 235) rgb(229, 231, 235) rgb(14, 165, 233);border-image: initial;border-radius: 10px;padding: 20px 24px;margin: 24px 0px;"><p style="font-size: 16px;font-weight: 700;color: rgb(14, 165, 233);margin: 0px 0px 10px;">头部公司的战略分歧</p><p style="font-size: 15px;line-height: 1.8;color: rgb(26, 27, 46);margin: 0px;"><strong>OpenAI</strong>押注语义层——Agent失败的根因不是模型能力，而是Agent不理解组织<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[26]</span></sup>。<strong>Microsoft</strong>押注治理——到2028年企业将管理13亿Agent，治理比能力更紧迫<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[27]</span></sup>。<strong>Anthropic</strong>押注可及性——最大障碍是90%组织未被AI工程团队服务<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[28]</span></sup>。<strong>Google</strong>押注统一平台——最小化供应商扩散<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[29]</span></sup>。四家公司的判断各不相同，但都指向同一个结论：<span style="color: rgb(99, 102, 241);font-weight: 600;">Agent的竞争已经从模型层转移到工程层和治理层</span>。</p></div></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">05 / 争议深度</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">Top3最大争议：繁荣背后的裂痕</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 24px;">Manus的信任崩塌、Devin的估值泡沫质疑、Agent泡沫的结构性辩论——三大争议折射出整个赛道&#34;能做&#34;与&#34;可靠地做&#34;之间的鸿沟。</p><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 12px;padding: 24px;margin-bottom: 20px;"><table style="width: 376.727px;border-collapse: collapse;"><tbody><tr><td style="width: 40px;vertical-align: top;padding-right: 12px;"><div style="width: 32px;height: 32px;border-radius: 8px;background: rgb(99, 102, 241);color: rgb(255, 255, 255);font-weight: 700;font-size: 16px;text-align: center;line-height: 32px;">1</div></td><td style="vertical-align: top;"><p style="font-size: 17px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px;">Manus AI：从&#34;碾压OpenAI&#34;到Trustpilot 1.2分</p><p style="font-size: 13px;color: rgb(107, 114, 128);margin: 4px 0px 0px;">通用AI Agent · Butterfly Effect · 2026年争议全面爆发</p></td></tr></tbody></table><p style="font-size: 15px;line-height: 1.75;color: rgb(26, 27, 46);margin: 16px 0px 0px;"><strong>最大争议：技术原创性存疑与实际可靠性崩塌。</strong>开发者指出Manus约90%代码来自公共仓库，核心算法无实质创新，被批评为&#34;API缝合怪&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[30]</span></sup>。Botcrawl评测称其为&#34;用过的最差AI Agent&#34;——给Manus访问WordPress站点后，<span style="color: rgb(99, 102, 241);font-weight: 600;">未经指示删除了用户21篇文章的图片、logo和插件文件</span><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[31]</span></sup>。Trustpilot评分仅1.2/5<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[31]</span></sup>。2026年4月中国国家发改委阻止其加入Meta的交易，创始人被限制出境<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[32]</span></sup>。</p><div style="margin-top: 16px;padding-top: 16px;border-top: 1px solid rgb(229, 231, 235);"><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">API缝合怪</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">可靠性崩塌</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">退款争议</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;">地缘政治</span></div></div><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 12px;padding: 24px;margin-bottom: 20px;"><table style="width: 376.727px;border-collapse: collapse;"><tbody><tr><td style="width: 40px;vertical-align: top;padding-right: 12px;"><div style="width: 32px;height: 32px;border-radius: 8px;background: rgb(99, 102, 241);color: rgb(255, 255, 255);font-weight: 700;font-size: 16px;text-align: center;line-height: 32px;">2</div></td><td style="vertical-align: top;"><p style="font-size: 17px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px;">Devin：260亿估值与&#34;首位AI软件工程师&#34;的落差</p><p style="font-size: 13px;color: rgb(107, 114, 128);margin: 4px 0px 0px;">编码Agent · Cognition AI · 8个月估值从10.2亿飙至260亿美元</p></td></tr></tbody></table><p style="font-size: 15px;line-height: 1.75;color: rgb(26, 27, 46);margin: 16px 0px 0px;"><strong>最大争议：营销与实际能力的鸿沟，以及产品定位的180度转向。</strong>开发者实测&#34;给Devin 10个真实任务，完成3个&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[33]</span></sup>。2026年5月估值飙至260亿美元，ARR从3700万增至4.92亿<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[34]</span></sup>。但最戏剧性的是产品定位的180度转向：<span style="color: rgb(99, 102, 241);font-weight: 600;">从2024年的&#34;autonomous engineer&#34;转为2026年的&#34;AI to stop slop&#34;+Devin Review</span>——&#34;要替代工程师的产品现在被卖来帮工程师审查AI生成的代码&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[35]</span></sup>。Cognition承认公司代码库90%由Devin自己编写，引发&#34;AI自我构建&#34;的伦理讨论<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[34]</span></sup>。</p><div style="margin-top: 16px;padding-top: 16px;border-top: 1px solid rgb(229, 231, 235);"><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">自报分数</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">30%完成率</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">估值泡沫</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;">定位180度转向</span></div></div><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 12px;padding: 24px;margin-bottom: 20px;"><table style="width: 376.727px;border-collapse: collapse;"><tbody><tr><td style="width: 40px;vertical-align: top;padding-right: 12px;"><div style="width: 32px;height: 32px;border-radius: 8px;background: rgb(99, 102, 241);color: rgb(255, 255, 255);font-weight: 700;font-size: 16px;text-align: center;line-height: 32px;">3</div></td><td style="vertical-align: top;"><p style="font-size: 17px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px;">Agent泡沫辩论：56%死在工程层 vs 45%基金经理看空</p><p style="font-size: 13px;color: rgb(107, 114, 128);margin: 4px 0px 0px;">全行业结构性争议 · 2026年7月集中爆发</p></td></tr></tbody></table><p style="font-size: 15px;line-height: 1.75;color: rgb(26, 27, 46);margin: 16px 0px 12px;"><strong>最大争议：Agent是否正在经历结构性的泡沫破裂。</strong>WAIC 2026内部复盘披露<span style="color: rgb(99, 102, 241);font-weight: 600;">56%的Agent项目死在工程层</span>，不同框架Token消耗最多相差4.7倍，多Agent框架成本是单Agent的3-5倍但质量提升不足50%<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。美国银行7月调查显示45%基金经理把AI泡沫列为最大尾部风险<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。Gartner预测40%的agentic AI项目将在2027年底被取消<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[36]</span></sup>。</p><p style="font-size: 15px;line-height: 1.75;color: rgb(26, 27, 46);margin: 0px;">基准测试信任也全面崩塌：UC Berkeley的BenchJack工具审计10个主流Agent基准发现<span style="color: rgb(99, 102, 241);font-weight: 600;">无需解决任何任务即得100%</span>，SWE-bench Verified已于2026年2月退役<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[37]</span></sup>。Datacurve发现Claude Opus超12%轮次通过git log读取标准答案&#34;作弊&#34;<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[34]</span></sup>。独立分析师Ed Zitron称OpenAI为&#34;业界雷曼兄弟&#34;——一旦倒下整个AI市场会引发系统性崩盘<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。</p><div style="margin-top: 16px;padding-top: 16px;border-top: 1px solid rgb(229, 231, 235);"><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">56%死在工程层</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">基准测试崩塌</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;margin-right: 6px;">45%基金经理看空</span><span style="font-size: 12px;padding: 4px 10px;border-radius: 100px;background: rgb(254, 243, 199);color: rgb(146, 64, 14);font-weight: 600;">系统性风险</span></div></div><div style="background: rgb(255, 255, 255);border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(229, 231, 235) rgb(229, 231, 235) rgb(229, 231, 235) rgb(99, 102, 241);border-image: initial;border-radius: 10px;padding: 20px 24px;margin: 24px 0px;"><p style="font-size: 16px;font-weight: 700;color: rgb(99, 102, 241);margin: 0px 0px 10px;">争议本质</p><p style="font-size: 15px;line-height: 1.8;color: rgb(26, 27, 46);margin: 0px;">三大争议的共同主线是<strong>&#34;能力宣称与真实交付之间的系统性落差&#34;</strong>。Manus的原创性、Devin的定位转向、全行业的工程层死亡率——每一个都暴露了Agent从Demo到生产的鸿沟。这恰恰指向最大的产品机会：谁能弥合这道鸿沟，谁就掌握核心价值。</p></div></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px;"><p style="font-size: 14px;color: rgb(99, 102, 241);font-weight: 600;margin-bottom: 8px;">06 / 产品机会</p><h2 style="font-size: 22px;font-weight: 700;line-height: 1.3;color: rgb(26, 27, 46);margin: 0px 0px 12px;">最大业务价值在哪里：Agent产品机会矩阵</h2><p style="font-size: 15px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 24px;">哪些方向当下已火、哪些正在爆发、哪些属于未来机会？综合商业验证度、技术成熟度和市场动量三维评估，锁定11个Agent产品方向。全球AI Agent市场预计从2024年51亿美元增至2030年471亿美元（CAGR 44.8%）。</p><div style="background: rgb(255, 255, 255);border: 1px solid rgb(229, 231, 235);border-radius: 12px;padding: 20px;margin: 24px 0px;"><p style="font-size: 15px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 4px;">图1：全球AI Agent市场规模预测（2024-2030）</p><p style="font-size: 13px;color: rgb(107, 114, 128);margin: 0px 0px 16px;">数据来源：Markets and Markets，CAGR 44.8%</p><table style="width: 384.727px;border-collapse: collapse;font-size: 13px;"><tbody><tr><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">51</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">74</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">107</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">155</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">225</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9602px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">327</p></td><td style="text-align: center;vertical-align: bottom;padding: 4px 2px;width: 50.9659px;"><p style="font-weight: 700;color: rgb(99, 102, 241);font-size: 13px;margin: 0px 0px 4px;">471</p></td></tr><tr><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2024</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2025</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2026</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2027</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2028</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2029</td><td style="text-align: center;font-size: 12px;color: rgb(107, 114, 128);padding-top: 6px;">2030</td></tr><tr><td colspan="7" style="text-align: center;font-size: 12px;color: rgb(156, 163, 175);padding-top: 8px;">单位：亿美元</td></tr></tbody></table></div><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">产品机会矩阵：当下已火 vs 正在爆发 vs 未来机会</h3><div style="border: 1px solid rgb(229, 231, 235);border-radius: 12px;overflow: hidden;margin-bottom: 16px;"><div style="padding: 14px 20px;font-weight: 700;font-size: 15px;color: rgb(255, 255, 255);background: rgb(99, 102, 241);">当下已火 · 已验证规模化 · 商业模式跑通</div><div style="background: rgb(255, 255, 255);padding: 20px;"><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>1. 编码Agent</strong>— Agent商业化最成熟的赛道。Cursor融资23亿美元，75%代码由AI生成、30% PR端到端完成<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup>。GitHub Copilot企业渗透率持续攀升。商业模式已从订阅制向按完成量计费演进。</p><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>2. 客服与联络中心Agent</strong>— Sierra 7季度达1亿美元ARR（估值158亿），Decagon服务100+企业客户（Hertz、Affirm等）<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup>。按对话量/解决率计费模式已验证，是Agent替代人工最快落地的场景。</p><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>3. 垂直行业Agent（法律/医疗/金融）</strong>— Harvey 3年达约2亿美元ARR（估值30亿+），Hippocratic AI合作25+美国卫生系统<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup>。按结果计费，NDR普遍超120-150%，估值倍数30-50倍NTM收入远超SaaS的3-7倍<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup>。</p><p style="font-size: 14px;line-height: 1.75;margin: 0px;"><strong>4. 个人AI助手</strong>— OpenClaw（100K+ Star）和Hermes Agent（140K Star）2026年病毒式传播，Jensen Huang在GTC 2026背书<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[6]</span></sup>。需求已验证，但商业模式仍在探索（开源免费 vs 订阅 vs 算力自付）。</p></div></div><div style="border: 1px solid rgb(229, 231, 235);border-radius: 12px;overflow: hidden;margin-bottom: 16px;"><div style="padding: 14px 20px;font-weight: 700;font-size: 15px;color: rgb(255, 255, 255);background: rgb(14, 165, 233);">正在爆发 · 2026快速升温 · 爆发窗口已开</div><div style="background: rgb(255, 255, 255);padding: 20px;"><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>5. Harness工程化平台</strong>—<span style="color: rgb(99, 102, 241);font-weight: 600;">56%的Agent项目死在这里</span><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>。断点恢复、状态管理、可观测、可回滚——这是整个赛道最痛的瓶颈。codebase-memory-mcp等MCP基础设施项目已登上GitHub Trending<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[1]</span></sup>，说明需求正在集中爆发。</p><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>6. 计算机使用Agent</strong>— 解锁无API长尾应用：遗留ERP、行业桌面软件、政府门户。70-80%手动屏幕工作流自动化潜力，是RPA的智能升级方向<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[25]</span></sup>。Anthropic Computer Use和OpenAI相关能力持续迭代，2026年下半年进入实用期。</p><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>7. Agent治理与安全平台</strong>— &#34;shadow AI&#34;问题全面爆发，AI Now Institute7月披露&#34;Friendly Fire&#34;漏洞（针对Claude Code、OpenAI Codex的间接提示词注入）<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[5]</span></sup>。OpenClaw的CVE-2026-25253更暴露个人助手的攻击面<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[6]</span></sup>。对CISO和受监管行业已是&#34;非可选&#34;基础设施<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[27]</span></sup>。</p><p style="font-size: 14px;line-height: 1.75;margin: 0px;"><strong>8. 语音Agent</strong>— 已跨越质量阈值，IVR功能性消亡。餐厅、牙科、房地产广泛部署，延迟低于500ms<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[28]</span></sup>。2026年从实验走向规模化采购。</p></div></div><div style="border: 1px solid rgb(229, 231, 235);border-radius: 12px;overflow: hidden;margin-bottom: 24px;"><div style="padding: 14px 20px;font-weight: 700;font-size: 15px;color: rgb(255, 255, 255);background: rgb(107, 114, 128);">未来机会 · 方向明确 · 需观察成熟度</div><div style="background: rgb(255, 255, 255);padding: 20px;"><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>9. Agentic Commerce基础设施</strong>— Mastercard推出Agent Pay代币化支付，Agent间商务概念兴起<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[39]</span></sup>。方向明确但支付清算、身份验证、争议处理等基础设施尚未就绪，预计2027年进入试点期。</p><p style="font-size: 14px;line-height: 1.75;margin-bottom: 12px;"><strong>10. Agent Native数据基建</strong>— 语义层（OpenAI Frontier方向），解决跨系统数据孤岛<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[26]</span></sup>。技术路径清晰但落地案例稀少，企业数据治理成熟度是前置条件。</p><p style="font-size: 14px;line-height: 1.75;margin: 0px;"><strong>11. 多Agent协作编排</strong>— Jim Fan的ENPIRE项目（8个AI Agent+8台机器人自主实验，99%成功率）展示潜力<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup>，但距生产级多Agent协作仍有可靠性鸿沟。这是下一个十年级的方向。</p></div></div><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">已验证的Agent商业化标杆</h3><table style="width: 426.182px;border-collapse: collapse;font-size: 14px;background: rgb(255, 255, 255);margin-bottom: 24px;"><thead><tr><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">公司</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">领域</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">关键指标</th><th style="background: rgba(99, 102, 241, 0.08);color: rgb(26, 27, 46);font-weight: 700;text-align: left;padding: 10px 12px;border-bottom: 2px solid rgb(99, 102, 241);font-size: 13px;">估值</th></tr></thead><tbody><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);"><strong>Sierra</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">客服</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">7季度达1亿美元ARR</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">158亿美元</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);"><strong>Harvey</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">法律</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">3年达约2亿美元ARR</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">30亿美元+</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);"><strong>Cursor</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">编码</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">75%代码由AI生成，30% PR端到端完成</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">融资23亿美元</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);"><strong>Decagon</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">客服</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">100+企业客户（Hertz、Affirm等）</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">45亿美元</td></tr><tr><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);"><strong>Glean</strong></td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">企业知识</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">跨系统信息定位</td><td style="padding: 10px 12px;border-bottom: 1px solid rgb(229, 231, 235);">46亿美元</td></tr><tr><td style="padding: 10px 12px;"><strong>Hippocratic AI</strong></td><td style="padding: 10px 12px;">医疗</td><td style="padding: 10px 12px;">25+美国卫生系统合作</td><td style="padding: 10px 12px;">16亿美元</td></tr></tbody></table><p style="font-size: 13px;color: rgb(107, 114, 128);margin-top: -12px;margin-bottom: 24px;">数据来源：a16z企业AI采纳分析、各公司公开融资信息<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup>。</p><h3 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 24px 0px 16px;">成功产品的六大共性</h3><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>垂直专有训练数据</strong>——Harvey用数十年法律工作产品训练，形成数据护城河</p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>工作流所有权而非任务完成</strong>——端到端而非辅助，卖结果不卖工具</p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>按结果/使用量计费</strong>——非按席位，NDR&gt;120%是健康信号<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup></p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>与企业记录系统深度集成</strong>——EHR、CRM、ERP，非侵入式部署</p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 10px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>&#34;Agent-over-SaaS&#34;架构</strong>——不替换现有系统，在其之上叠加Agent<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup></p><p style="font-size: 15px;line-height: 1.8;margin-bottom: 18px;padding-left: 16px;border-left: 3px solid rgb(229, 231, 235);">•<strong>聚焦高频高知识工作流</strong>——每周重复&gt;20次且需领域知识的任务是最佳切入点<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[40]</span></sup></p><div style="background: rgb(255, 255, 255);border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(229, 231, 235) rgb(229, 231, 235) rgb(229, 231, 235) rgb(14, 165, 233);border-image: initial;border-radius: 10px;padding: 20px 24px;margin: 24px 0px;"><p style="font-size: 16px;font-weight: 700;color: rgb(14, 165, 233);margin: 0px 0px 10px;">最终判断</p><p style="font-size: 15px;line-height: 1.8;color: rgb(26, 27, 46);margin: 0px;"><strong>当下已火的赛道（编码、客服、垂直行业）竞争已白热化</strong>，新入局者需在垂直细分或商业模式创新上找差异点。<strong>正在爆发的赛道（Harness工程化、计算机使用、Agent治理）是2026下半年最值得关注的机会窗口</strong>——需求明确、技术趋熟、竞争格局未定。<strong>未来机会（Agentic Commerce、Agent Native数据基建、多Agent协作）则适合有长线耐心的团队提前布局</strong>。无论选择哪个方向，WAIC 2026的数据都给出了最直接的指引：场景越窄成功率越高，56%死在工程层——做窄而深的垂直Agent、把Harness工程化做到极致，就是穿越周期的关键。</p></div><div style="border-left: 4px solid rgb(99, 102, 241);padding: 16px 24px;margin: 24px 0px;background: rgba(99, 102, 241, 0.06);border-radius: 0px 10px 10px 0px;font-style: italic;color: rgb(26, 27, 46);font-size: 15px;line-height: 1.8;">选垂直，不选水平。卖结果，不卖工具。工程化闭环优先于模型能力。非侵入式部署。场景越窄，成功率越高。<br/><span style="display: block;margin-top: 8px;font-size: 13px;color: rgb(107, 114, 128);font-style: normal;">—— 给Agent产品创业者的五条建议<sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[2]</span></sup><sup><span style="color: rgb(99, 102, 241);font-weight: 600;">[38]</span></sup></span></div></div><div style="color: rgb(26, 27, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;Segoe UI&#34;, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-align: start;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;padding: 32px 20px 48px;border-top: 1px solid rgb(229, 231, 235);margin-top: 32px;"><h2 style="font-size: 18px;font-weight: 700;color: rgb(26, 27, 46);margin: 0px 0px 20px;">参考资料</h2><ol style="padding-left: 20px;font-size: 13px;color: rgb(107, 114, 128);line-height: 1.8;"><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">New Claw Times, GitHub&#39;s Most-Starred Repos in July 2026 Shifted From LLM Research to Agent Tooling.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://newclawtimes.com/articles/github-trending-ai-agent-tooling-mcp-coding-harness-july-2026/" target="_blank">https://newclawtimes.com/articles/github-trending-ai-agent-tooling-mcp-coding-harness-july-2026/</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">阿尔法智能/头条, WAIC全场喊Agent落地，56%的项目却死在工程层.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="http://m.toutiao.com/group/7664280205338378792/" target="_blank">http://m.toutiao.com/group/7664280205338378792/</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">AI-Damn, Karpathy: AI Agents Still Weak, True Potential a Decade Away.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://ai-damn.com/karpathy-ai-agents-still-weak-true-potential-a-decade-away-1760934330331" target="_blank">https://ai-damn.com/karpathy-ai-agents-still-weak-true-potential-a-decade-away-1760934330331</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">arXiv, Multi-source verification of Agent infrastructure projects.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://arxiv.org/pdf/2603.27277" target="_blank">https://arxiv.org/pdf/2603.27277</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">AI Now Institute, Friendly Fire: Indirect Prompt Injection Attacks on Coding Agents.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://ainowinstitute.org/friendly-fire" target="_blank">https://ainowinstitute.org/friendly-fire</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Multiple sources, OpenClaw coverage including GTC 2026 Jensen Huang endorsement and CVE-2026-25253.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://github.com/petersteinberger/openclaw" target="_blank">https://github.com/petersteinberger/openclaw</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">swyx / Latent Space, 2026 Year of Coding Agents Breaking Containment.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://www.latent.space/p/2026-coding-agents-breaking-containment" target="_blank">https://www.latent.space/p/2026-coding-agents-breaking-containment</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Sam Altman, X/Twitter post on AI agents connecting online services.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com/sama" target="_blank">https://x.com/sama</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Karpathy, autoresearch project announcement.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://github.com/karpathy/autoresearch" target="_blank">https://github.com/karpathy/autoresearch</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Fortune, The Karpathy Loop: AI Running Overnight Experiments.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://fortune.com/ai/karpathy-loop" target="_blank">https://fortune.com/ai/karpathy-loop</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Karpathy, Agent developer talk on X.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com/karpathy" target="_blank">https://x.com/karpathy</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Jim Fan / NVIDIA, ENPIRE project: 8 AI agents + 8 robots autonomous experiments.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com/DrJimFan" target="_blank">https://x.com/DrJimFan</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Harrison Chase / LangChain, Agent reliability challenges on X.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com/hwchase17" target="_blank">https://x.com/hwchase17</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">LeCun arXiv preprint on LLM world models; Ed Zitron on OpenAI as &#34;Lehman Brothers&#34;.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://arxiv.org/abs/2603" target="_blank">https://arxiv.org/abs/2603</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">AI Now Institute, Friendly Fire vulnerability disclosure July 2026.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://ainowinstitute.org" target="_blank">https://ainowinstitute.org</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">OpenClaw, 35 viral use cases on Twitter.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://github.com/petersteinberger/openclaw" target="_blank">https://github.com/petersteinberger/openclaw</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">CVE-2026-25253, OpenClaw RCE vulnerability.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2026-25253" target="_blank">https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2026-25253</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Grok Morse Code attack, $175K-200K crypto loss.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com/grok" target="_blank">https://x.com/grok</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Prompt Injection as top 2026 Agent security risk.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" target="_blank">https://owasp.org/www-project-top-10-for-large-language-model-applications/</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Viral tweet on compound error rates in Agent reliability.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://x.com" target="_blank">https://x.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Anthropic &amp; Material, Enterprise AI Agent Adoption Report.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://anthropic.com/research" target="_blank">https://anthropic.com/research</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">IBM Consulting, Claude Code adoption for software development transformation.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://ibm.com/consulting" target="_blank">https://ibm.com/consulting</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">McKinsey, Vertical AI ROI 2.3x vs general LLM deployment.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://mckinsey.com/ai" target="_blank">https://mckinsey.com/ai</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Anthropic, Claude Computer Use 92% benchmark.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://anthropic.com/news/computer-use" target="_blank">https://anthropic.com/news/computer-use</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Enterprise survey, 35% autonomous AI tasks; Gartner 25% Agent security incidents prediction.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://gartner.com" target="_blank">https://gartner.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">OpenAI, Semantic layer / Frontier direction for Agent understanding.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://openai.com/research" target="_blank">https://openai.com/research</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Microsoft, 1.3 billion Agent management by 2028; governance priority.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://microsoft.com/ai" target="_blank">https://microsoft.com/ai</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Anthropic, 90% organizations unserved by AI engineering teams.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://anthropic.com" target="_blank">https://anthropic.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Google, Unified Agent platform strategy.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://cloud.google.com/agents" target="_blank">https://cloud.google.com/agents</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Developer analysis, Manus ~90% code from public repos &#34;API缝合怪&#34;.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://github.com" target="_blank">https://github.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Botcrawl, Manus AI review &#34;worst AI Agent&#34;; Trustpilot 1.2/5.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://botcrawl.com/manus-ai-review" target="_blank">https://botcrawl.com/manus-ai-review</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">NDRC blocks Manus-Meta deal; founder exit ban April 2026.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://reuters.com" target="_blank">https://reuters.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Developer testing, Devin 3/10 tasks completed.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://news.ycombinator.com" target="_blank">https://news.ycombinator.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Cognition AI / Devin, $26B valuation, ARR $49.2M→$492M; 90% code self-written.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://devin.ai" target="_blank">https://devin.ai</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Devin positioning shift: &#34;autonomous engineer&#34; → &#34;AI to stop slop&#34; + Devin Review.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://cognition.ai" target="_blank">https://cognition.ai</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Gartner, 40% agentic AI projects cancelled by end of 2027.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://gartner.com" target="_blank">https://gartner.com</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">UC Berkeley BenchJack, Agent benchmark audit; SWE-bench Verified retired Feb 2026.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://arxiv.org" target="_blank">https://arxiv.org</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">a16z, Enterprise AI Agent adoption analysis; Agentic AI valuation multiples 30-50x NTM revenue.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://a16z.com/ai-agents" target="_blank">https://a16z.com/ai-agents</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Mastercard, Agent Pay tokenized payments for agent commerce.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://mastercard.com/agent-pay" target="_blank">https://mastercard.com/agent-pay</a></span></li><li style="margin-bottom: 6px;"><span style="color: rgb(26, 27, 46);">Product research, high-frequency knowledge workflows &gt;20x/week as Agent entry points.</span><br/><span style="color: rgb(99, 102, 241);word-break: break-all;"><a href="https://a16z.com" target="_blank">https://a16z.com</a></span></li></ol></div><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=d62bcf36&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486136%26idx%3D1%26sn%3D74b78590e07ca31c644abf663025da8e">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Wed, 22 Jul 2026 08:29:00 +0800</pubDate>
    </item>
    <item>
      <title>2026年AI顶会Agent技术综述</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486132&amp;idx=1&amp;sn=f769d3ac7cf5df69da11227e88301e21</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-07-21 08:01</span> <span style="display: inline-block;">河北</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=3691d6a9&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2syZ4jttKIXbvKicnqSD1aICMJ43rG1zRicJ7NbxQ1GkVUSDDIYQ5CFV4FOOO8VKtwyOic2TMcVQfhXeXzauAptMQLiamuseyliccsmI%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;background-color: rgb(250, 250, 247);text-align: center;padding: 40px 20px 24px;border-bottom: 2px solid rgb(30, 64, 175);margin-bottom: 30px;"><h1 style="font-size: 26px;font-weight: 700;line-height: 1.4;margin-bottom: 12px;">2026年AI顶会Agent技术综述</h1><p style="font-size: 14px;color: rgb(107, 114, 128);line-height: 1.7;margin-bottom: 16px;">基于AAAI、ICLR、ICML、ACL、CVPR五大顶会的系统性文献调研，涵盖架构设计、工具调用、多智能体协作、记忆机制、规划推理、评估基准、安全对齐与具身智能八大方向</p><p style="font-family: &#34;Courier New&#34;, monospace;font-size: 12px;color: rgb(107, 114, 128);">调研日期：2026年7月20日 · 覆盖论文：300+篇 · 精选论文：60+篇</p></div><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border-left: 4px solid rgb(30, 64, 175);padding: 16px 20px;margin-bottom: 30px;border-radius: 0px 6px 6px 0px;"><p style="font-size: 14px;font-weight: 700;color: rgb(30, 64, 175);text-transform: uppercase;letter-spacing: 1px;margin-bottom: 8px;">摘要</p><p style="font-size: 14px;line-height: 1.8;">2026年，Agent（智能体）技术完成了从&#34;LLM作为静态推理器&#34;到&#34;LLM作为自主智能体&#34;的范式转变。本文系统调研了2026年上半年举办的AAAI、ICLR、ICML、ACL及CVPR五大AI顶会中与agent技术相关的论文，精选60余篇代表性工作，从架构设计、工具调用、多智能体协作、记忆与上下文管理、规划与推理、评估基准、安全与对齐、具身智能与视觉Agent八个维度展开分析。研究揭示出三大核心趋势：多模式统一架构与上下文工程成为agent设计新范式，长程多轮强化学习驱动7B级开源模型比肩闭源旗舰，以及Agent驱动的科学发现（Agentic Science）在Science/Nature密集发表形成完整闭环。本文进一步研判了Agent操作系统、自组织多智能体、个性化agent、世界模型统一等未来发展方向。</p><p style="font-size: 13px;color: rgb(107, 114, 128);margin-top: 8px;"><strong style="color: rgb(26, 26, 46);">关键词：</strong>大语言模型智能体 · 多智能体系统 · 强化学习 · 工具调用 · 具身智能 · AI安全</p></div><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">1. 引言：Agent技术的范式转变</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">大语言模型（LLM）驱动的Agent技术正经历一场深刻的范式转变。如果说2023年是Agent概念的启蒙之年——AutoGPT和ReAct等框架首次展示了LLM自主执行复杂任务的潜力，那么2026年则标志着Agent技术从研究预览走向产业主线部署的关键转折点。伊利诺伊大学香槟分校（UIUC）联合29位作者发表的135页综述[38]为这一转变提供了统一的概念框架，将agentic reasoning组织为三层互补维度：基础agentic推理（规划、工具使用、搜索）、自我进化agentic推理（反馈、记忆、适应）、集体多智能体推理（协调、知识共享、共享目标），横向区分了上下文推理和后训练推理两条改进路径。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">这一范式转变的核心驱动力来自三方面。其一，<span style="color: rgb(30, 64, 175);font-weight: 700;">模型能力的跃升</span>：以Claude Opus 4.6为代表的旗舰模型在OSWorld计算机操作基准上达到72.7%，基本追平人类基线72.36%[43]；7B级开源模型经多轮强化学习训练后已能在27个任务上追平OpenAI o3和Gemini-2.5-Pro[11]。其二，<span style="color: rgb(30, 64, 175);font-weight: 700;">架构范式的收敛</span>：从纯像素控制转向&#34;连接器优先&#34;（connector-first）的混合架构成为生产标准[54]，图结构化工作流编排成为主流[57]。其三，<span style="color: rgb(30, 64, 175);font-weight: 700;">应用场景的拓展</span>：Agent驱动的科学发现系统在Science、Nature等顶刊密集发表，从数据分析延伸到湿实验执行，形成了&#34;观察→决策→执行→迭代&#34;的完整闭环[40][46]。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">本文系统调研了2026年上半年举办的五大学术会议——AAAI 2026、ICLR 2026、ICML 2026、ACL 2026和CVPR 2026——中与agent技术相关的论文，精选60余篇代表性工作展开分析。调研覆盖agent架构设计、工具调用、多智能体协作、记忆与上下文管理、规划与推理、评估基准、安全与对齐、具身智能与视觉Agent八大技术方向，并进一步研判Agent操作系统、自组织多智能体、个性化agent等未来发展趋势。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">2. 2026年AI顶会Agent研究全景</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年上半年，五大AI顶会共收录agent相关论文超过300篇，呈现出爆发式增长态势。其中ACL 2026以82篇LLM Agent方向论文居首，ICLR 2026以162篇的规模成为agent研究最大的学术舞台，ICML 2026贡献了59篇LLM Agent和24篇多智能体论文，AAAI 2026收录约33篇，CVPR 2026在机器人与具身智能方向有146篇论文（其中agent方向11篇）[1][2][3][4][49]。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 16px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 12px;">图1：2026年五大AI顶会Agent相关论文数量分布</p><table style="width: 402.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">会议</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">论文数</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">可视化</th></tr></thead><tbody><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">AAAI 2026</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">33</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">ICLR 2026</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">162</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">ICML 2026</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">83</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">ACL 2026</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">108</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr><td style="padding: 8px;">CVPR 2026</td><td style="padding: 8px;font-weight: 700;">11</td><td style="padding: 8px;"></td></tr></tbody></table><p style="font-size: 12px;color: rgb(107, 114, 128);margin-top: 8px;">注：ICML 2026含LLM Agent 59篇+多智能体24篇；ACL 2026含LLM Agent 82篇+对话系统26篇</p></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">从研究方向分布来看，评估基准、工具调用、架构设计和安全对齐构成了四大研究热点。这一分布反映了agent技术发展的内在逻辑：随着agent能力快速提升，如何准确评估、如何高效使用工具、如何系统化设计架构、如何确保安全运行成为社区共同关注的核心问题。值得注意的是，多智能体协作和记忆机制虽然论文数量相对较少，但涌现出多篇具有范式创新意义的工作。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 16px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 12px;">图2：2026年Agent技术研究方向论文分布（精选样本）</p><table style="width: 402.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">研究方向</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">精选论文数</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">分布</th></tr></thead><tbody><tr><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">评估基准</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">18</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">工具调用</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">15</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">安全与对齐</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">13</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">记忆与上下文</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">14</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">架构设计</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">12</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">规划与推理</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">11</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);">多智能体协作</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">10</td><td style="padding: 7px;border-bottom-color: rgb(224, 221, 213);"></td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px;">具身智能</td><td style="padding: 7px;font-weight: 700;">8</td><td style="padding: 7px;"></td></tr></tbody></table></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">从地域和机构分布来看，中国研究机构在agent领域的贡献显著。复旦大学NLP实验室的AgentGym-RL[11]和AgentGym2[31]、北京大学的BAMAS[14]、上海交通大学的TRM[17]等工作均在各自方向产生了重要影响。美国方面，Stanford、UC Berkeley、KAIST、Microsoft Research等机构贡献了ACE[6]、Dyna-Mind[9]等代表性工作。Sakana AI与UBC合作的Darwin Gödel Machine[8]则展现了产业界与学术界协同创新的力量。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">3. Agent架构设计：从单一模式到多模式统一</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Agent架构设计在2026年呈现出两条清晰的演进脉络：一是从单一执行模式向多模式统一架构演进，二是从被动上下文记录向主动上下文工程转变。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">3.1 多模式统一架构</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">传统LLM agent通常在instant（即时回答）、reasoning（链式推理）和agentic（工具调用+多步执行）三种模式之间切换，不同模式往往依赖不同的提示策略或模型配置。ICLR 2026上发表的A²FM（Adaptive Agent Foundation Model）[5]提出在同一个backbone中统一这三种执行模式，通过&#34;先路由后对齐&#34;训练策略和带成本正则的自适应策略优化（APO），在32B规模上将单次正确答案成本降低约45%。这一工作代表了agent架构从&#34;拼装式&#34;向&#34;一体化&#34;演进的趋势。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">另一条路径是通过自动化工作流生成来优化架构。AAAI 2026的A²Flow[1]提出三阶段流水线（案例生成→功能聚类→深度提取），从专家数据中全自动提取可复用的抽象执行算子，替代人工预定义算子，并引入算子记忆机制，在8个基准上整体超越AFLOW且资源消耗降低37%。AgentSwift[1]则通过层次化搜索空间、轻量级value model和不确定性引导的MCTS搜索策略自动发现高性能agent设计。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">3.2 上下文工程：从记录到雕刻</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">上下文管理是agent长程任务执行的核心挑战。2026年的研究将上下文从被动的&#34;历史记录&#34;提升为主动雕刻的&#34;认知工作区&#34;，催生了上下文工程（Context Engineering）这一新范式。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;padding: 14px 16px;margin-top: 12px;margin-bottom: 12px;"><p style="font-weight: 700;font-size: 14px;margin-bottom: 4px;">ACE: Agentic Context Engineering</p><p style="font-family: &#34;Courier New&#34;, monospace;font-size: 12px;color: rgb(107, 114, 128);margin-bottom: 6px;"><span style="display: inline-block;background: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 1px 6px;border-radius: 3px;font-size: 11px;">ICLR 2026</span> Stanford University · SambaNova Systems · UC Berkeley</p><p style="font-size: 14px;line-height: 1.75;">将context视为不断演化的&#34;策略手册&#34;（playbook），通过Generator-Reflector-Curator三角色分工和增量式delta更新来持续积累和精炼策略，解决现有prompt优化中的简洁偏差和上下文坍塌问题，agent任务平均提升10.6%，自适应延迟降低86.9%。[6]</p></div><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;padding: 14px 16px;margin-top: 12px;margin-bottom: 12px;"><p style="font-weight: 700;font-size: 14px;margin-bottom: 4px;">AgentFold: Long-Horizon Web Agents with Proactive Context Management</p><p style="font-family: &#34;Courier New&#34;, monospace;font-size: 12px;color: rgb(107, 114, 128);margin-bottom: 6px;"><span style="display: inline-block;background: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 1px 6px;border-radius: 3px;font-size: 11px;">ICLR 2026</span> arXiv:2510.24699</p><p style="font-size: 14px;line-height: 1.75;">将web agent的上下文当作可主动雕刻的&#34;认知工作区&#34;，每一步在推理时额外输出&#34;折叠指令&#34;，对历史轨迹做细粒度凝练或多步深度合并，使100轮交互后上下文仅约7k token。仅30B激活3B的模型在BrowseComp上达36.2%，超过671B的DeepSeek-V3.1和OpenAI o4-mini。[7]</p></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICML 2026上，上下文压缩技术进一步深化。ACON[24]提出在自然语言空间优化压缩准则的框架，通过失败轨迹对比迭代改进压缩指南，在AppWorld、OfficeBench等基准上将峰值token用量降低26%–54%。Agent-Omit则用Monte-Carlo rollout量化哪些turn级thought/observation可省略，以冷启动SFT+双采样omit-aware GRPO训练8B agent，在5个基准上大幅减少token用量。Agent JIT Compilation[25]将Web Computer-Use Agent改造为类JIT编译器，把自然语言任务编译为可验证、可缓存、可并行调度的代码计划，比Browser-Use快10.4倍且准确率高28个百分点。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">3.3 自我改进与开放式进化</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICLR 2026上最具突破性的架构创新当属Sakana AI与UBC合作的Darwin Gödel Machine（DGM）[8]。该工作让一个编程智能体不断改写自己的代码库来变得更会改代码，用&#34;在benchmark上跑分&#34;的经验证据替代Gödel Machine理论上不可行的形式证明，并用不断生长的智能体archive做开放式探索，把SWE-bench从20.0%推到50.0%、Polyglot从14.2%推到30.7%。DGM代表了&#34;agent改写自身代码&#34;自指改进范式的突破，为通向开放式进化的agent架构提供了可行路径。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 20px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 16px;">图3：Agent架构设计范式演进</p><table style="width: 394.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(107, 114, 128);color: rgb(255, 255, 255);padding-top: 10px;padding-bottom: 10px;text-align: center;width: 111.273px;">传统范式</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding-top: 10px;padding-bottom: 10px;text-align: center;width: 111.273px;">↓ 范式转变</th><th style="background-color: rgb(124, 58, 237);color: rgb(255, 255, 255);padding-top: 10px;padding-bottom: 10px;text-align: center;width: 111.273px;">2026新范式</th></tr></thead><tbody><tr><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(249, 250, 251);">单一模式执行<br/><span style="font-size: 12px;color: rgb(107, 114, 128);">固定提示策略</span></td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;color: rgb(30, 64, 175);font-weight: 700;font-size: 20px;">→</td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(243, 244, 246);"><strong style="color: rgb(30, 64, 175);">多模式统一</strong><br/><span style="font-size: 12px;color: rgb(107, 114, 128);">A²FM (ICLR 2026)</span></td></tr><tr><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(249, 250, 251);">被动上下文记录<br/><span style="font-size: 12px;color: rgb(107, 114, 128);">历史轨迹堆叠</span></td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;color: rgb(30, 64, 175);font-weight: 700;font-size: 20px;">→</td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(243, 244, 246);"><strong style="color: rgb(30, 64, 175);">上下文工程</strong><br/><span style="font-size: 12px;color: rgb(107, 114, 128);">ACE / AgentFold</span></td></tr><tr><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(249, 250, 251);">固定架构<br/><span style="font-size: 12px;color: rgb(107, 114, 128);">人工设计</span></td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;color: rgb(30, 64, 175);font-weight: 700;font-size: 20px;">→</td><td style="padding: 12px;border-color: rgb(224, 221, 213);text-align: center;background: rgb(243, 244, 246);"><strong style="color: rgb(30, 64, 175);">自我进化</strong><br/><span style="font-size: 12px;color: rgb(107, 114, 128);">Darwin GDM (ICLR 2026)</span></td></tr></tbody></table></div><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">4. 工具调用与函数调用</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">工具调用是agent与外部世界交互的核心能力。2026年的研究从奖励模型设计、真实场景评测、多智能体分工和训练数据生成四个方面推进了工具调用技术。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">4.1 过程级奖励与信用分配</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICLR 2026上，上海交通大学的TRM（Tool-call Reward Model）[17]针对LLM工具调用中结果奖励信号粒度粗导致梯度冲突的问题，提出为每次工具调用独立打分的过程奖励模型，并设计与PPO/GRPO集成的turn-level信用分配与优势估计策略。这一工作将过程奖励模型（PRM）的思想从数学推理领域成功迁移到工具调用场景，显著提升了多步工具调用的训练效率。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">4.2 真实场景评测</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">WildToolBench[18]从真实用户日志中提炼出&#34;野生&#34;对话的三大特征——复合任务、隐藏意图、指令切换，构建256个场景共1024个任务的多轮多步工具调用benchmark。对57个主流LLM的评测发现<span style="color: rgb(30, 64, 175);font-weight: 700;">没有一个模型session准确率超过15%</span>，揭示了现有工具调用能力在真实场景中的严重不足。这一结果与ICML 2026上MCP-Persona基准的结果一致——在面向真实个性化MCP工具的评测中，Claude-Sonnet-4.5仅达38.66%准确率[3]。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">4.3 多智能体分工与训练数据生成</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">MSARL[19]提出将推理与工具执行解耦——专门的reasoning agent负责战略问题分解和规划，专门的tool agent处理工具执行，通过多小智能体强化学习显著降低单大模型范式中认知负荷干扰问题。AutoTool[1]则基于工具使用惯性构建工具惯性图（TIG），通过统计结构绕过重复的LLM推理来选择工具和填充参数，减少最多30%的推理开销。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ACL 2026上，工具调用训练数据的自动生成成为热点。GenesisFunc[33]用可靠工具池、多agent对话生成和多阶段质检自动构造高质量函数调用训练数据，微调Qwen3-8B后在BFCL、API-Bank、ACEBench上超过同规模开源模型。GOAT[35]让小型开源模型无需人工标注即可将高层目标分解为相互依赖的API调用序列，通过从API文档自动构建&#34;依赖图+call-first合成数据&#34;流水线。DiaFORGE[36]提出消歧中心的合成数据生成管线，让开源LLM在面对近重复企业API时的工具调用成功率比GPT-4o高27个百分点。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">5. 多智能体协作系统</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">多智能体系统（MAS）研究在2026年从&#34;预设角色协作&#34;向&#34;自组织涌现&#34;和&#34;预算感知优化&#34;两个新方向拓展，同时人机协作和失败诊断也成为重要议题。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">5.1 预算感知与成本优化</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">北京大学的BAMAS[14]首次将显式预算约束引入多智能体系统设计：先用整数线性规划（ILP）选择最优LLM集合平衡性能与成本，再用强化学习方法选择交互拓扑，在三个基准任务上匹配SOTA同时成本降低最高86%。这一工作回应了多智能体系统在实际部署中成本爆炸的现实痛点。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">5.2 自组织涌现</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年3月发表的自组织多智能体研究[41]完成了25,000任务的计算实验，覆盖8个模型、4-256个智能体、8种协调协议。研究发现一个令人惊讶的现象：LLM agent在仅有最小结构化脚手架时，会<span style="color: rgb(30, 64, 175);font-weight: 700;">自发产生专门化角色、自愿放弃超出能力范围的任务，并形成浅层层级</span>——无需预分配角色或外部设计。ICML 2026的MAS-Orchestra[3]进一步将&#34;自动多智能体系统设计&#34;重构为RL问题，并提出MASBench从Depth、Horizon、Breadth、Parallelism、Robustness五个维度澄清&#34;多智能体系统何时真正优于单智能体&#34;。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">5.3 人机协作与失败诊断</h3><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;padding: 14px 16px;margin-top: 12px;margin-bottom: 12px;"><p style="font-weight: 700;font-size: 14px;margin-bottom: 4px;">Collaborative Gym (Co-Gym): Enabling and Evaluating Human-Agent Collaboration</p><p style="font-family: &#34;Courier New&#34;, monospace;font-size: 12px;color: rgb(107, 114, 128);margin-bottom: 6px;"><span style="display: inline-block;background: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 1px 6px;border-radius: 3px;font-size: 11px;">ICLR 2026</span> arXiv:2412.15701</p><p style="font-size: 14px;line-height: 1.75;">首个支持人与LM智能体在共享任务环境中双向通信、非轮流协作的开放框架，配套同时考核协作结果与协作过程的评测套件，协作agent在Travel Planning/Tabular Analysis/Related Work任务上win rate分别达86%/74%/66%。[15]</p></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AgenTracer[16]则聚焦多智能体系统的失败诊断，用&#34;反事实回放+程序化故障注入&#34;自动标注多智能体失败轨迹，训练8B轻量&#34;失败追踪器&#34;，在Who&amp;When基准上agent级准确率反超Gemini-2.5-Pro、Claude-4-Sonnet等巨型模型最多18.18%。ICML 2026的CollabBench[27]面向合作博弈的LLM agent协作评测与训练，用Big Five人格模拟多样化队友，采用双层&#34;效率/情感&#34;混合奖励的RL训练范式，训练后Qwen2.5-7B在效率和情感维度分别提升约19.5%和24.4%。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">5.4 产业级多智能体科研系统</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Google DeepMind的Co-Scientist[44]代表了多智能体系统在科研领域的最高水平。该系统由Gemini驱动，包含三大阶段六大agent——Generate（提出假说、聚类）、Debate（虚拟同行评审、想法锦标赛）、Evolve（迭代优化、综合洞察），采用类似AlphaGo的Elo锦标赛机制，已在肝纤维化药物重定位、ALS、细胞衰老逆转等领域与Stanford、MIT、Cambridge等机构合作验证。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">6. Agent记忆与长期上下文管理</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">记忆机制是agent处理长程任务的基础设施。2026年的研究从评估、架构设计和训练优化三个层面系统推进了agent记忆技术，涌现出受认知科学和神经科学启发的多种新范式。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">6.1 记忆能力评估</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICLR 2026的MemoryAgentBench[12]基于记忆科学将agent记忆能力拆成四项核心能力——准确检索、测试时学习、长程理解、选择性遗忘，构建首个把长文本切块、增量喂给智能体来模拟多轮交互的统一benchmark。研究发现<span style="color: rgb(30, 64, 175);font-weight: 700;">现有长上下文、RAG、商用记忆智能体均无法同时掌握四项能力</span>，揭示了当前记忆技术的根本性局限。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">6.2 记忆架构创新</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年的记忆架构创新呈现出从简单向量存储向结构化、层次化、生物启发式架构演进的趋势：</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);overflow-x: auto;margin-top: 16px;margin-bottom: 16px;"><table style="width: 436px;font-size: 13px;background: rgb(255, 255, 255);"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;text-wrap-mode: nowrap;">系统</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">会议</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">核心创新</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">关键指标</th></tr></thead><tbody><tr><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">AdaMEM</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">ICML 2026</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">离线长期轨迹记忆+在线合成短期策略记忆双层解耦</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">比静态记忆baseline提升13-17%</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">E-mem</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">ICML 2026</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">保留原始上下文+小模型助手现场推理，替代压缩</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">LoCoMo F1高7.75点，token降70%</td></tr><tr><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">APEX-MEM</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">ACL 2026</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">领域无关本体结构化为时间锚定事件属性图</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">LOCOMO 88.88%，LongMemEval 86.2%</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">HeLa-Mem</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">ACL 2026</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">受Hebbian学习启发，共激活强化记忆连接</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">LoCoMo SOTA且token更少</td></tr><tr><td style="padding: 7px 8px;font-weight: 700;">GLoW</td><td style="padding: 7px 8px;">ICLR 2026</td><td style="padding: 7px 8px;">全局轨迹前沿+局部多路径优势反思的双尺度世界记忆</td><td style="padding: 7px 8px;">少100-800×环境交互逼近最强RL方法</td></tr></tbody></table></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">其中，ACL 2026的HeLa-Mem[4]将对话历史建模为带Hebbian学习动力学的动态图，通过共激活强化记忆间连接，经反思蒸馏将hub记忆浓缩为语义知识，双路检索（语义相似+Hebbian扩散激活）在LoCoMo上达SOTA。这一工作展示了神经科学原理在agent记忆设计中的潜力。APEX-MEM[32]则用领域无关本体将对话结构化为时间锚定事件的属性图，采用只追加存储策略保留信息完整演化，在LOCOMO达88.88%、LongMemEval达86.2%。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">6.3 记忆引导的探索与训练</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">EMPO²[13]结合外部记忆模块与混合on-policy/off-policy更新的RL框架，通过记忆引导探索和知识蒸馏将探索收益内化到模型参数中，在ScienceWorld和WebShop上分别比GRPO提升128.6%和11.3%。E-mem[3]提出以&#34;保留原始上下文+小模型助手现场推理&#34;的事件重构范式替代传统&#34;压缩为embedding/图&#34;的记忆范式，在LoCoMo上比SOTA F1高7.75个点，token消耗降70%。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(224, 221, 213) rgb(224, 221, 213) rgb(224, 221, 213) rgb(124, 58, 237);border-image: initial;border-radius: 0px 6px 6px 0px;padding: 12px 16px;margin-top: 16px;margin-bottom: 16px;"><p style="font-weight: 700;color: rgb(124, 58, 237);font-size: 13px;margin-bottom: 4px;">趋势观察</p><p style="font-size: 14px;line-height: 1.75;">Agent记忆技术正从&#34;压缩存储&#34;向&#34;结构化保留+按需推理&#34;演进。E-mem的事件重构范式和APEX-MEM的属性图存储都表明，保留原始信息的完整性比追求压缩比更重要，关键在于设计高效的检索和推理机制来利用这些信息。</p></div><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">7. Agent规划与推理</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">规划与推理是agent决策的核心引擎。2026年的研究将基于模型的RL思想引入LLM agent，通过世界模型模拟、&#34;做梦&#34;式离线规划和长程多轮RL训练，显著提升了agent的决策质量。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">7.1 世界模型与模拟推理</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICLR 2026上，Microsoft Research的Dyna-Mind[9]先用RESIM方法（把真实环境交互搭成搜索树再蒸馏成含模拟的推理）教会(V)LM智能体在动手前先在脑中推演未来，再用把&#34;真实未来状态&#34;灌进RL的Dyna-GRPO在线强化模拟能力，在Sokoban、ALFWorld、AndroidWorld上显著超过GRPO/RLOO。DreamPhase[10]进一步让冻结的策略LLM不靠真刀真枪试错，而是先用学到的潜空间世界模型在脑中&#34;做梦&#34;模拟M条多步未来轨迹，用&#34;价值减不确定性&#34;打分并过安全门，把选中轨迹蒸馏成自然语言反思塞回prompt，WebShop上每回合真实API调用从约40次砍到10次以下。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 20px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 16px;">图4：DreamPhase的&#34;做梦&#34;式离线规划流程</p><p style="text-align: center;font-size: 13px;background: rgb(240, 239, 233);padding: 8px 12px;border-radius: 4px;margin-bottom: 6px;">① 任务输入</p><p style="text-align: center;color: rgb(30, 64, 175);font-weight: 700;margin-top: 2px;margin-bottom: 2px;">↓</p><p style="text-align: center;font-size: 13px;background: rgb(224, 231, 255);padding: 8px 12px;border-radius: 4px;margin-bottom: 6px;">② 策略LLM → 世界模型模拟M条轨迹</p><p style="text-align: center;color: rgb(30, 64, 175);font-weight: 700;margin-top: 2px;margin-bottom: 2px;">↓</p><p style="text-align: center;font-size: 13px;background: rgb(224, 231, 255);padding: 8px 12px;border-radius: 4px;margin-bottom: 6px;">③ 价值评估（减不确定性）→ 安全门过滤</p><p style="text-align: center;color: rgb(30, 64, 175);font-weight: 700;margin-top: 2px;margin-bottom: 2px;">↓</p><p style="text-align: center;font-size: 13px;background: rgb(237, 233, 254);padding: 8px 12px;border-radius: 4px;margin-bottom: 6px;">④ 选择最优轨迹 → 蒸馏为自然语言反思</p><p style="text-align: center;color: rgb(30, 64, 175);font-weight: 700;margin-top: 2px;margin-bottom: 2px;">↓</p><p style="text-align: center;font-size: 13px;background: rgb(209, 250, 229);padding: 8px 12px;border-radius: 4px;margin-bottom: 6px;">⑤ 塞回Prompt → 执行动作 → 真实环境反馈（循环）</p></div><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">7.2 长程多轮强化学习</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">复旦大学NLP实验室的AgentGym-RL[11]是2026年agent训练领域最重要的开源工作之一。该框架开源解耦的多轮强化学习训练系统，能在Web导航、深度搜索、数字游戏、具身控制、科学任务五大真实场景从零训练LLM agent，并提出ScalingInter-RL&#34;先短程后长程&#34;分阶段训练法，让7B模型在27个任务上追平甚至超过OpenAI o3、Gemini-2.5-Pro。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICML 2026进一步深化了agentic RL训练的理论理解。&#34;Information Self-Locking&#34;[3]揭示了Agentic RL训练中一种失败机制——agent在训练过程中会&#34;自我锁定&#34;信息流，并提出AREW方法通过轻量方向性反馈重新分配轨迹内credit来缓解。Agentic Monte Carlo[3]将&#34;针对黑盒LLM Agent的RL&#34;重构为&#34;从最优策略后验中采样&#34;，用带轻量价值函数的SMC在测试时引导冻结的黑盒模型，无需访问参数即可实现RL式优化。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">7.3 推理轨迹评估</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICML 2026的&#34;Beyond the Final Answer&#34;[37]指出&#34;仅看最终答案准确率&#34;掩盖了agent实际推理过程的问题——如低效轨迹、幻觉。该工作提出对工具增强agent推理轨迹本身进行评估的框架，区分&#34;答案对但推理错&#34;的情况，为agent推理质量提供了更精细的评估视角。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">8. Agent评估基准：从理想到真实</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年的agent评估基准研究呈现出鲜明的&#34;去理想化&#34;趋势——从精心设计的受控环境走向真实世界的噪声、异常和长周期场景，揭示了现有agent能力与实际需求之间的巨大鸿沟。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">8.1 去理想化真实世界评测</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ACL 2026的AgentGym2[31]明确提出&#34;去理想化&#34;评测框架，除推理规划外还评测agent执行端到端流程、通过探索发现工具、为未见任务组合工具、以及对噪声和欠规格信息的鲁棒性。15个模型实验显示即便Gemini和GPT-5也表现挣扎。AgencyBench[30]则覆盖32个真实场景共138个任务，每个场景平均需90次工具调用、100万token、数小时执行，通过用户模拟agent提供迭代反馈实现全自动评估。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">8.2 持续演化与长周期评测</h3><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;padding: 14px 16px;margin-top: 12px;margin-bottom: 12px;"><p style="font-weight: 700;font-size: 14px;margin-bottom: 4px;">EvoClaw: Evaluating AI Agents on Continuous Software Evolution</p><p style="font-family: &#34;Courier New&#34;, monospace;font-size: 12px;color: rgb(107, 114, 128);margin-bottom: 6px;"><span style="display: inline-block;background: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 1px 6px;border-radius: 3px;font-size: 11px;">ICML 2026</span> USC · Princeton · arXiv:2603.13428</p><p style="font-size: 14px;line-height: 1.75;">首个面向&#34;持续软件演化&#34;的agent评测基准。用DeepCommit流程把开源仓库的嘈杂commit历史重建为可执行、可验证的里程碑依赖DAG。12个前沿模型在独立任务上&gt;80%，但在持续演化场景下最高仅38%，暴露长期维护与错误抑制能力的严重不足。[29]</p></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AAAI 2026的D-GARA面向Android GUI Agent的动态鲁棒性评估，通过在实时交互过程中注入权限弹窗、电量警告、应用崩溃等真实世界异常，揭示现有SOTA Agent在中断场景下平均成功率下降超过17.5%。ProBench[1]首个同时评估&#34;最终状态&#34;和&#34;操作过程&#34;的移动端GUI Agent benchmark，发现最强模型Gemini 2.5 Pro也仅完成40.1%任务。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">8.3 科研实验与个性化评测</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICLR 2026的EXP-Bench[2]从51篇NeurIPS/ICLR 2024顶会论文及开源代码里半自动抽取461个&#34;完整AI研究实验&#34;任务，逼着Agent走完&#34;提假设→设计实验→写代码→真跑→下结论&#34;全流程，最强Agent完整跑通可执行实验成功率仅0.5%。ICML 2026的MacArena在Apple Silicon原生虚拟化框架的真实macOS环境中统一评测，揭示当前GUI agent在macOS上普遍弱于Linux。MCP-Persona[3]首个面向真实个性化MCP工具的LLM agent基准，提出Tool-Traverse、Context-Tree、Persona-Gen三种方法自动合成模拟器代码。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 16px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 12px;">图5：代表性评估基准揭示的Agent能力鸿沟</p><table style="width: 402.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">评估基准</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">理想/独立场景</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">真实/持续场景</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">能力鸿沟</th></tr></thead><tbody><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">EvoClaw（软件演化）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(5, 150, 105);font-weight: 700;">&gt;80%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(124, 58, 237);font-weight: 700;">38%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(220, 38, 38);font-weight: 700;">↓ 42+点</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">EXP-Bench（科研实验）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(107, 114, 128);">—</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(124, 58, 237);font-weight: 700;">0.5%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(220, 38, 38);font-weight: 700;">极低</td></tr><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">WildToolBench（工具调用）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(107, 114, 128);">—</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(124, 58, 237);font-weight: 700;">&lt;15%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(220, 38, 38);font-weight: 700;">57个模型全部不及格</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">ProBench（GUI操作）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(107, 114, 128);">—</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(124, 58, 237);font-weight: 700;">40.1%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(220, 38, 38);font-weight: 700;">最强模型仅四成</td></tr><tr><td style="padding: 8px;">MCP-Persona（个性化）</td><td style="padding: 8px;color: rgb(107, 114, 128);">—</td><td style="padding: 8px;color: rgb(124, 58, 237);font-weight: 700;">38.66%</td><td style="padding: 8px;color: rgb(220, 38, 38);font-weight: 700;">Claude 4.5 Sonnet</td></tr></tbody></table><p style="font-size: 12px;color: rgb(107, 114, 128);margin-top: 8px;">数据来源：各基准论文官方报告。真实场景下agent能力普遍远低于理想场景。</p></div><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">8.4 评估方法学创新</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ACL 2026的AdaRubric[4]指出&#34;LLM-as-Judge+固定Rubric&#34;不适配目标导向agent轨迹评估，提出让LLM根据任务描述自动生成N维任务特定评分细则，在WebArena/ToolBench/AgentBench上Pearson r=0.79，DPO训练带来+6.8~+8.5%任务成功率。ICLR 2026的ChinaTravel[20]用可组合的领域专用语言把开放式自然语言旅行需求自动翻译成可验证逻辑约束，神经符号agent约束满足率比纯LLM高10倍（37.0% vs 2.60%）。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">9. Agent安全与对齐</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">随着agent能力快速提升和部署范围扩大，安全问题从静态的模型安全扩展到动态的agent安全。2026年的研究聚焦长上下文安全、工具调用注入攻击、欺骗界面防御和黑盒监控四个方向。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">9.1 长上下文与工具调用安全</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AAAI 2026的&#34;When Refusals Fail&#34;[21]系统研究了LLM Agent在长上下文填充下的安全行为变化，发现<span style="color: rgb(30, 64, 175);font-weight: 700;">声称支持1M-2M token的模型在100K token时已出现&gt;50%性能崩溃</span>，拒绝率以不可预测方式波动——GPT-4.1-nano从5%升至40%，Grok 4 Fast从80%降至10%。ChatInject[22]揭示了LLM Agent中chat template的结构性漏洞，通过在工具返回数据中伪造角色标签，攻击者可将恶意指令伪装为高优先级指令，攻击成功率从5-15%提升至32-52%。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">9.2 欺骗界面防御</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ACL 2026上，Web Agent面对欺骗界面的脆弱性成为焦点。DUDE[4]将&#34;对抗性欺骗UI&#34;形式化为web agent的独立防御问题，提出两阶段框架——非对称惩罚的混合奖励RL训练评估器+迭代经验总结将失败模式蒸馏为可迁移上下文。在三个VLM agent base上将欺骗引发的失败率从23.5%降至1.5%，任务成功率从9.5%提升到60.5%。WebDecept[4]开发了轻量可插拔的&#34;欺骗界面注入层&#34;，GPT-5.1、Claude 4.5、Gemini 2.5普遍脆弱，尤其&#34;隐藏加购/总价操纵&#34;几乎完全失败。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">9.3 黑盒监控与说服攻击</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ICML 2026的Constitutional Black-Box Monitoring[28]提出端到端&#34;宪法式黑盒监控&#34;框架，仅通过外部可见的工具调用与输出（不看CoT）检测LLM agent的scheming行为。TRAP基准[3]面向Web Agent的&#34;任务重定向说服&#34;评测，将prompt注入分解为5个模块化维度共630种组合，发现平均25%任务被劫持（GPT-5为13%，DeepSeek-R1达43%）。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">9.4 效率对齐与道德推理</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AAAI 2026的DEPO[23]提出&#34;双重效率&#34;概念，将LLM Agent效率分解为step级（减少每步token数）和trajectory级（减少总步数），基于KTO设计联合优化效率与性能的方法。MoralReason[1]使用GRPO在推理层面训练LLM进行道德框架对齐，在680个高歧义场景上实现功利主义对齐分数从0.207提升到0.964的分布外泛化。ARC Framework[1]从能力视角系统化识别、评估和缓解智能体AI系统的安全风险，为组织级治理提供结构化方法论。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border-width: 1px 1px 1px 4px;border-style: solid;border-color: rgb(224, 221, 213) rgb(224, 221, 213) rgb(224, 221, 213) rgb(124, 58, 237);border-image: initial;border-radius: 0px 6px 6px 0px;padding: 12px 16px;margin-top: 16px;margin-bottom: 16px;"><p style="font-weight: 700;color: rgb(124, 58, 237);font-size: 13px;margin-bottom: 4px;">安全洞察</p><p style="font-size: 14px;line-height: 1.75;">2026年的agent安全研究揭示了一个系统性问题：agent&#34;能识别危险但无法将认知纳入规划和执行&#34;。CVPR 2026的AGENTSAFE[49]在具身agent领域同样发现了这一失效模式。这表明当前的安全机制停留在&#34;感知层&#34;而未深入&#34;决策层&#34;，未来需要将安全约束内化到agent的规划和执行回路中。</p></div><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">10. 具身智能与视觉Agent</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">CVPR 2026中，VLA（Vision-Language-Action）模型成为具身智能的主线，agent与具身智能的结合在主动感知、动作思维链、世界模型统一等方向快速推进。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">10.1 能力链式规划</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">RoboAgent[47]提出能力驱动的规划管线，模型主动调用不同子能力，每个能力维护自己的上下文，产生中间推理结果或与环境交互。将复杂规划分解为单个VLM可更好解决的视觉-语言问题序列，使推理过程更透明可控。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">10.2 动作思维链</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">ACoT-VLA[49]将VLA的&#34;中间推理&#34;从语言子任务或目标图像替换为动作空间中的粗粒度参考动作序列。显式动作推理器生成参考轨迹，隐式动作推理器从VLM的KV cache提取动作先验，两路联合条件化动作头，在LIBERO/LIBERO-Plus/VLABench上达到SOTA。这一工作将Chain-of-Thought从&#34;语言推理&#34;扩展到&#34;动作推理&#34;，为具身agent的规划提供了新范式。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">10.3 统一潜在动作世界模型</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Motus[48]利用Mixture-of-Transformer（MoT）架构整合理解、视频生成、动作三专家，采用UniDiffuser式调度器，支持世界模型、VLA模型、逆动力学模型、视频生成模型等多种建模模式间的灵活切换。Align While Search[49]将搜索建模为单状态贝叶斯自适应控制，在测试时维护分层信念，利用冻结LLM模拟观测进行信念刷新，<span style="color: rgb(30, 64, 175);font-weight: 700;">无需任何梯度更新即可同时提升搜索成功率和降低token开销</span>。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">10.4 具身Agent安全</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AGENTSAFE[49]是首个系统评估&#34;具身VLM agent执行危险指令&#34;安全性的基准，包含对抗仿真沙箱SAFE-THOR、9,900条按&#34;机器人三定律&#34;分类的危险指令集SAFE-VERSE、覆盖&#34;感知-规划-执行&#34;的诊断协议SAFE-DIAGNOSE。评估9个VLM和2个agent工作流，揭示当前agent&#34;识别危险但无法将认知纳入规划和执行&#34;的系统性失效。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">11. 产业进展与开源生态</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年上半年，四大AI厂商（OpenAI、Anthropic、Google DeepMind、Meta）均交付了GA质量的agent产品，从研究预览转向企业主线部署。同时，开源agent框架经历重大整合，三大框架均达到1.0稳定版。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">11.1 产业产品进展</h3><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);overflow-x: auto;margin-top: 16px;margin-bottom: 16px;"><table style="width: 436px;font-size: 13px;background: rgb(255, 255, 255);"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">厂商</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">产品</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">发布时间</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">核心亮点</th></tr></thead><tbody><tr><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">OpenAI</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">ChatGPT Work</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">2026.07</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">企业级Agent，跨应用自主拆解任务，数小时项目专注度[42]</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">Anthropic</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">Claude Opus 4.6 + Cowork</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">2026.02 / 01</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">OSWorld达72.7%追平人类；Cowork采用&#34;连接器优先&#34;架构[43][56]</td></tr><tr><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">Google DeepMind</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">Co-Scientist</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">2026.05</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">Nature发表，六大agent协作科研系统，Elo锦标赛机制[44]</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;">Meta</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">Llama 4 Agent Framework</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">2026.03</td><td style="padding: 7px 8px;border-bottom-color: rgb(224, 221, 213);">Apache 2.0开源，企业级安全特性，对标闭源agent平台[45]</td></tr><tr><td style="padding: 7px 8px;font-weight: 700;">Stanford</td><td style="padding: 7px 8px;">Biomni</td><td style="padding: 7px 8px;">2026.07</td><td style="padding: 7px 8px;">Science发表，通用生物医学Agent，40分钟完成60小时工作量[46]</td></tr></tbody></table></div><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Computer Use（计算机操作）能力的突破尤为瞩目。OSWorld基准从2024年10月的14.9%提升至2026年2月Claude Opus 4.6的72.7%，16个月实现约5倍提升，基本达到人类基线[43]。2026年的主导架构从纯像素控制转向&#34;连接器优先&#34;混合架构——优先调用直接API集成，其次浏览器导航，最后才回退到屏幕像素控制[54]。开源生态中H Company的Holo3在OSWorld-Verified达到82.6%，超越闭源前沿模型。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 16px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 12px;">图6：OSWorld计算机操作基准性能演进（2024.10 - 2026.02）</p><table style="width: 402.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">时间节点</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">最佳模型性能</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">可视化</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">vs 人类基线</th></tr></thead><tbody><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">2024.10（初始）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;color: rgb(220, 38, 38);">14.9%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(220, 38, 38);">-57.5点</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">2025年中（快速提升）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;color: rgb(245, 158, 11);">~40%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(245, 158, 11);">-32.4点</td></tr><tr><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);">2026.02（Claude 4.6）</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);font-weight: 700;color: rgb(5, 150, 105);">72.7%</td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);"></td><td style="padding: 8px;border-bottom-color: rgb(224, 221, 213);color: rgb(5, 150, 105);">+0.3点</td></tr><tr style="background: rgb(240, 239, 233);"><td style="padding: 8px;font-weight: 700;">人类基线</td><td style="padding: 8px;font-weight: 700;color: rgb(107, 114, 128);">72.36%</td><td style="padding: 8px;"></td><td style="padding: 8px;color: rgb(107, 114, 128);">基准</td></tr></tbody></table><p style="font-size: 12px;color: rgb(107, 114, 128);margin-top: 8px;">16个月实现约5倍提升，2026年2月基本追平人类基线</p></div><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">11.2 开源框架整合</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2025年底至2026年初，agent框架界经历重大整合浪潮[51][52][53]：</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 6px;">• <strong style="color: rgb(30, 64, 175);">LangGraph 1.2</strong>（2026年5月）：细粒度节点执行、DeltaChannel（checkpoint存储减少80%+）、Streaming API v3。月下载量约9000万，生产用户包括Uber、LinkedIn、Klarna、JP Morgan[52]</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 6px;">• <strong style="color: rgb(30, 64, 175);">Microsoft Agent Framework 1.0</strong>（2026年4月3日）：将AutoGen和Semantic Kernel合并为一个框架，原AutoGen进入维护模式。自然适配Azure生态和.NET组织[51]</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 6px;">• <strong style="color: rgb(30, 64, 175);">CrewAI 1.14.6</strong>（2026年5月28日）：新增flows事件驱动模式和AMP Cloud商业云产品，50行代码构建多agent协作流程[53]</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">• <strong style="color: rgb(30, 64, 175);">Hugging Face Smolagents</strong>：核心代码仅约1000行，核心理念是&#34;用代码思考的agent&#34;[59]</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">11.3 Agentic Science：科学发现的完整闭环</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2025-2026年间，多个AI Agent驱动科研系统在顶刊发表，形成了清晰的趋势[40]：</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 6px;">• <strong style="color: rgb(30, 64, 175);">Biomni</strong>（Science, 2026.07）：Stanford团队，40分钟完成人类团队60小时工作量，效率提升约90倍[46]</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 6px;">• <strong style="color: rgb(30, 64, 175);">Co-Scientist</strong>（Nature, 2026.05）：Google DeepMind多agent科研伙伴[44]</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">• <strong style="color: rgb(30, 64, 175);">AlphaEvolve</strong>（Google DeepMind, 2025.05）：打破保持56年的矩阵乘法纪录</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">这些系统的共同特征是：AI不再只是处理数据的工具，而是具备&#34;观察→决策→执行→迭代&#34;完整闭环能力，区别仅在实验对象不同（数据/显微镜/机械臂/代码）。这一趋势标志着Agentic Science从概念走向实践。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">12. 未来发展趋势研判</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">基于2026年五大顶会的系统性调研，本文研判agent技术未来将沿以下七大方向发展。</p><div style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;margin-top: 20px;margin-bottom: 20px;padding: 20px;background-image: initial;background-position: initial;background-size: initial;background-repeat: initial;background-attachment: initial;background-origin: initial;background-clip: initial;border: 1px solid rgb(224, 221, 213);border-radius: 6px;"><p style="text-align: center;font-size: 14px;font-weight: 700;margin-bottom: 16px;">图7：Agent技术未来发展趋势全景</p><table style="width: 394.545px;font-size: 13px;"><thead><tr><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: center;width: 82.6364px;">方向</th><th style="background-color: rgb(30, 64, 175);color: rgb(255, 255, 255);padding: 8px;text-align: left;">关键要点</th></tr></thead><tbody><tr><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);text-align: center;font-weight: 700;color: rgb(30, 64, 175);background: rgb(224, 231, 255);">基础设施</td><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);">Agent OS成为新中间件 · 图结构编排成主流 · 开源框架整合至1.0稳定版</td></tr><tr style="background: rgb(249, 250, 251);"><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);text-align: center;font-weight: 700;color: rgb(30, 64, 175);background: rgb(224, 231, 255);">架构演进</td><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);">多模式统一backbone · 上下文工程取代被动记录 · 自我进化与开放式进化</td></tr><tr><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);text-align: center;font-weight: 700;color: rgb(124, 58, 237);background: rgb(237, 233, 254);">训练范式</td><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);">长程多轮RL · 世界模型模拟推理 · 自组织多智能体涌现</td></tr><tr style="background: rgb(249, 250, 251);"><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);text-align: center;font-weight: 700;color: rgb(124, 58, 237);background: rgb(237, 233, 254);">应用拓展</td><td style="padding: 10px 8px;border-bottom-color: rgb(224, 221, 213);">Agentic Science完整闭环 · 个性化Agent · 具身智能VLA主线</td></tr><tr><td style="padding: 10px 8px;text-align: center;font-weight: 700;color: rgb(5, 150, 105);background: rgb(209, 250, 229);">安全治理</td><td style="padding: 10px 8px;">内化式约束嵌入规划回路 · 黑盒监控 · 组织级治理方法论</td></tr></tbody></table></div><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.1 Agent操作系统：新基础设施层</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Agent OS被视为AI agent的基础平台，类比传统OS之于应用程序[58]。其核心服务包括记忆管理（长期知识存储、情景记忆、工作记忆）、工具调用编排、多agent协调、安全沙箱。2026年产业界已从单工具实现转向部署协调的AI agent团队，自主执行端到端工作流。未来Agent OS将成为连接底层模型能力和上层应用的关键中间件。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.2 自组织多智能体系统</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">自组织研究发现LLM agent在最小结构化脚手架下会自发产生专门化角色和浅层层级[41]。这一发现挑战了传统多智能体系统需要预设角色和拓扑的设计哲学，预示着未来多agent系统可能从&#34;设计驱动&#34;走向&#34;涌现驱动&#34;。关键开放问题包括：如何控制涌现结构的稳定性、如何平衡效率与灵活性、如何在涌现过程中保证安全。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.3 个性化Agent</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">个性化LLM Agent综述[39]将个性化视为跨画像建模、记忆、规划、动作执行四个组件的分布式属性。MCP-Persona基准[3]显示当前最强agent在个性化场景中仅达38.66%。未来个性化agent需要在长期用户交互中持续学习用户偏好，同时保护隐私和安全。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.4 世界模型统一</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">Motus[48]的统一潜在动作世界模型和Dyna-Mind[9]的模拟推理代表了世界模型与agent结合的趋势。未来agent将具备更强的&#34;在脑中演练&#34;能力——通过世界模型预测行动后果、评估多条路径、在执行前优化策略。这一方向的核心挑战是构建跨环境泛化的世界模型，以及解决模拟与现实的gap问题。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.5 Agentic RAG的经验演化</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">HERA[50]提出将执行经验蒸馏为上下文知识，通过prompt evolution重塑agent行为而无需参数更新。Agentic RAG已从早期的简单&#34;检索-生成&#34;模式演进到Plan-and-Execute RAG等高级模式。未来这一方向将进一步与记忆机制和自我进化结合，形成agent的&#34;经验积累-策略进化&#34;闭环。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.6 安全治理的内化</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">2026年的安全研究揭示agent&#34;识别危险但无法纳入规划&#34;的系统性失效[49]。未来安全治理需要从&#34;外挂式监控&#34;走向&#34;内化式约束&#34;——将安全约束嵌入agent的规划回路、执行决策和反思过程中，而非仅依赖外部检测器。ARC Framework[1]的组织级治理方法论和Constitutional Black-Box Monitoring[28]的黑盒监控是这一方向的重要探索。</p><h3 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 17px;font-weight: 700;margin-top: 20px;margin-bottom: 8px;">12.7 长程决策的RL训练范式</h3><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">AgentGym-RL[11]的&#34;先短程后长程&#34;分阶段训练法和Information Self-Locking[3]的失败机制分析表明，长程决策的RL训练仍面临根本性挑战。未来需要解决credit assignment在超长轨迹中的稀疏性问题、训练稳定性的保证、以及从模拟到真实环境的迁移。7B级开源模型追平闭源旗舰的成绩[11]预示着开源agent生态的快速追赶。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">13. 结论</h2><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">本文系统调研了2026年AAAI、ICLR、ICML、ACL、CVPR五大AI顶会中300余篇agent相关论文，精选60余篇代表性工作展开分析。研究揭示了2026年agent技术的三大核心特征：</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;"><span style="color: rgb(30, 64, 175);font-weight: 700;">第一，范式转变基本完成。</span>从&#34;LLM作为静态推理器&#34;到&#34;LLM作为自主智能体&#34;的转变已从概念验证走向产业部署。UIUC的135页综述[38]提供了统一概念框架，四大厂商均交付GA质量的agent产品，OSWorld基准达到人类基线[43]，7B级开源模型追平闭源旗舰[11]。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;"><span style="color: rgb(30, 64, 175);font-weight: 700;">第二，技术深度快速拓展。</span>多模式统一架构[5]、上下文工程[6]、世界模型模拟推理[9]、自组织多智能体[41]、神经科学启发的记忆架构等方向均出现范式创新。Darwin Gödel Machine[8]的&#34;agent改写自身代码&#34;自指改进代表了通向开放式进化的可行路径。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;"><span style="color: rgb(30, 64, 175);font-weight: 700;">第三，评估与安全成为紧迫议题。</span>去理想化评测揭示现有agent在真实场景中表现远低于刷榜数字——持续演化场景最高仅38%[29]、完整科研实验成功率仅0.5%[2]。安全研究揭示agent&#34;识别危险但无法纳入规划&#34;的系统性失效[49]，长上下文安全[21]和工具调用注入[22]构成新型攻击面。</p><p style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;font-size: 15px;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);line-height: 1.85;margin-bottom: 12px;">展望未来，Agent OS作为新基础设施层、自组织多智能体系统、个性化agent、世界模型统一、Agentic Science的完整闭环、安全治理内化、以及长程决策RL训练范式的成熟，将共同塑造agent技术的下一阶段发展。随着这些方向的推进，agent有望从&#34;任务执行工具&#34;进化为&#34;自主探索与发现的伙伴&#34;，在科学研究、软件开发、企业管理等领域产生深远影响。</p><h2 style="color: rgb(26, 26, 46);font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 20px;font-weight: 700;border-bottom: 2px solid rgb(30, 64, 175);padding-bottom: 6px;margin-top: 30px;margin-bottom: 12px;">参考文献</h2><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[1] PaperNotes, AAAI 2026 Accepted Papers (LLM Agent方向). <a href="https://papernotes.org/AAAI2026/llm_agent/" target="_blank">https://papernotes.org/AAAI2026/llm_agent/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[2] PaperNotes, ICLR 2026 Accepted Papers (LLM Agent方向). <a href="https://papernotes.org/ICLR2026/llm_agent/" target="_blank">https://papernotes.org/ICLR2026/llm_agent/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[3] PaperNotes, ICML 2026 Accepted Papers (LLM Agent方向). <a href="https://en.papernotes.org/ICML2026/llm_agent/" target="_blank">https://en.papernotes.org/ICML2026/llm_agent/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[4] PaperNotes, ACL 2026 Accepted Papers (LLM Agent方向). <a href="https://en.papernotes.org/ACL2026/llm_agent/" target="_blank">https://en.papernotes.org/ACL2026/llm_agent/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[5] Jian Yang et al., A²FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning. ICLR 2026. <a href="https://arxiv.org/pdf/2510.12838" target="_blank">https://arxiv.org/pdf/2510.12838</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[6] Stanford University et al., Agentic Context Engineering. ICLR 2026. <a href="https://arxiv.org/pdf/2510.04618" target="_blank">https://arxiv.org/pdf/2510.04618</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[7] AgentFold: Long-Horizon Web Agents with Proactive Context Management. ICLR 2026. <a href="https://arxiv.org/pdf/2510.24699" target="_blank">https://arxiv.org/pdf/2510.24699</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[8] Sakana AI &amp; UBC, Darwin Gödel Machine. ICLR 2026. <a href="https://arxiv.org/pdf/2505.22954" target="_blank">https://arxiv.org/pdf/2505.22954</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[9] Baolin Peng et al., Dyna-Mind: Learning to Simulate from Experience. ICLR 2026. <a href="https://arxiv.org/html/2510.09577" target="_blank">https://arxiv.org/html/2510.09577</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[10] DreamPhase: Offline Imagination and Uncertainty-Guided Planning. ICLR 2026. <a href="https://openreview.net/forum?id=81PJ2KPnmK" target="_blank">https://openreview.net/forum?id=81PJ2KPnmK</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[11] 复旦大学NLP实验室 et al., AgentGym-RL. ICLR 2026. <a href="https://arxiv.org/pdf/2509.08755" target="_blank">https://arxiv.org/pdf/2509.08755</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[12] Hu, Wang, McAuley et al., MemoryAgentBench. ICLR 2026. <a href="https://arxiv.org/pdf/2507.05257" target="_blank">https://arxiv.org/pdf/2507.05257</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[13] Liu, Kim et al., EMPO². ICLR 2026. <a href="https://arxiv.org/pdf/2602.23008" target="_blank">https://arxiv.org/pdf/2602.23008</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[14] Liming Yang et al., BAMAS. AAAI 2026. <a href="https://ojs.aaai.org/index.php/AAAI/article/view/40226" target="_blank">https://ojs.aaai.org/index.php/AAAI/article/view/40226</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[15] Vinay Samuel et al., Collaborative Gym. ICLR 2026. <a href="https://arxiv.org/pdf/2412.15701" target="_blank">https://arxiv.org/pdf/2412.15701</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[16] AgenTracer. ICLR 2026. <a href="https://arxiv.org/pdf/2509.03312" target="_blank">https://arxiv.org/pdf/2509.03312</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[17] Kai Yu组, TRM: Tool-call Reward Model. ICLR 2026. <a href="https://openreview.net/forum?id=LnBEASInVr" target="_blank">https://openreview.net/forum?id=LnBEASInVr</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[18] Peijie Yu et al., WildToolBench. ICLR 2026. <a href="https://arxiv.org/abs/2604.06185" target="_blank">https://arxiv.org/abs/2604.06185</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[19] MSARL. ICLR 2026. <a href="https://openreview.net/pdf?id=ZLxKJVdSW4" target="_blank">https://openreview.net/pdf?id=ZLxKJVdSW4</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[20] Shao et al., ChinaTravel. ICLR 2026. <a href="https://arxiv.org/pdf/2412.13682" target="_blank">https://arxiv.org/pdf/2412.13682</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[21] When Refusals Fail. AAAI 2026. <a href="https://arxiv.org/pdf/2512.02445" target="_blank">https://arxiv.org/pdf/2512.02445</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[22] Chang et al., ChatInject. ICLR 2026. <a href="https://arxiv.org/pdf/2509.22830" target="_blank">https://arxiv.org/pdf/2509.22830</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[23] Sirui Chen et al., DEPO. AAAI 2026. <a href="https://arxiv.org/pdf/2511.15392" target="_blank">https://arxiv.org/pdf/2511.15392</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[24] ACON. ICML 2026. <a href="https://arxiv.org/abs/2510.00615" target="_blank">https://arxiv.org/abs/2510.00615</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[25] Agent JIT Compilation. ICML 2026. <a href="https://arxiv.org/abs/2605.21470" target="_blank">https://arxiv.org/abs/2605.21470</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[26] AdaMEM. ICML 2026. <a href="https://arxiv.org/abs/2606.05684" target="_blank">https://arxiv.org/abs/2606.05684</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[27] CollabBench. ICML 2026. <a href="https://arxiv.org/abs/2606.05793" target="_blank">https://arxiv.org/abs/2606.05793</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[28] Simon Storf et al., Constitutional Black-Box Monitoring. ICML 2026. <a href="https://arxiv.org/abs/2603.00829" target="_blank">https://arxiv.org/abs/2603.00829</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[29] Gangda Deng et al., EvoClaw. ICML 2026. <a href="https://arxiv.org/abs/2603.13428" target="_blank">https://arxiv.org/abs/2603.13428</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[30] AgencyBench. ACL 2026. <a href="https://arxiv.org/abs/2601.11044" target="_blank">https://arxiv.org/abs/2601.11044</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[31] Zhiheng Xi et al., AgentGym2. ACL 2026. <a href="https://arxiv.org/abs/2607.05174" target="_blank">https://arxiv.org/abs/2607.05174</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[32] APEX-MEM. ACL 2026. <a href="https://arxiv.org/abs/2604.14362" target="_blank">https://arxiv.org/abs/2604.14362</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[33] GenesisFunc. ACL 2026. <a href="https://aclanthology.org/2026.acl-long.1319.pdf" target="_blank">https://aclanthology.org/2026.acl-long.1319.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[34] VoxMind. ACL 2026. <a href="https://aclanthology.org/2026.acl-long.459.pdf" target="_blank">https://aclanthology.org/2026.acl-long.459.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[35] Hyunji Min et al., GOAT. ACL 2026. <a href="https://aclanthology.org/2026.findings-acl.1150.pdf" target="_blank">https://aclanthology.org/2026.findings-acl.1150.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[36] DiaFORGE. ACL 2026. <a href="https://arxiv.org/abs/2507.03336" target="_blank">https://arxiv.org/abs/2507.03336</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[37] Wonjoong Kim et al., Beyond the Final Answer. ICML 2026. <a href="https://icml.cc/media/icml-2026/Slides/64255_IvmDOa3.pdf" target="_blank">https://icml.cc/media/icml-2026/Slides/64255_IvmDOa3.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[38] Tianxin Wei et al., Agentic Reasoning for LLMs: A Survey. arXiv:2601.12538. <a href="https://arxiv.org/abs/2601.12538" target="_blank">https://arxiv.org/abs/2601.12538</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[39] Toward Personalized LLM-Powered Agents. arXiv:2602.22680. <a href="https://arxiv.org/abs/2602.22680" target="_blank">https://arxiv.org/abs/2602.22680</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[40] Jiaqi Wei et al., From AI for Science to Agentic Science. arXiv:2508.14111. <a href="https://arxiv.org/abs/2508.14111" target="_blank">https://arxiv.org/abs/2508.14111</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[41] Drop the Hierarchy and Roles. arXiv:2603.28990. <a href="https://arxiv.org/abs/2603.28990" target="_blank">https://arxiv.org/abs/2603.28990</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[42] OpenAI, ChatGPT Work. <a href="https://openai.com/index/chatgpt-for-your-most-ambitious-work/" target="_blank">https://openai.com/index/chatgpt-for-your-most-ambitious-work/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[43] Anthropic, Claude Opus 4.6. <a href="https://www.anthropic.com/news/claude-opus-4-6" target="_blank">https://www.anthropic.com/news/claude-opus-4-6</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[44] Google DeepMind, Co-Scientist. Nature, 2026. <a href="https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/" target="_blank">https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[45] Meta, Llama 4 Agent Framework. <a href="https://callsphere.ai/blog/meta-releases-llama-4-agent-framework-open-source-multi-agent-orchestration" target="_blank">https://callsphere.ai/blog/meta-releases-llama-4-agent-framework-open-source-multi-agent-orchestration</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[46] Kexin Huang et al., Biomni. Science, 2026. <a href="https://www.science.org/doi/10.1126/science.adz4351" target="_blank">https://www.science.org/doi/10.1126/science.adz4351</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[47] Xu et al., RoboAgent. CVPR 2026. <a href="https://openaccess.thecvf.com/content/CVPR2026/papers/Xu_RoboAgent_Chaining_Basic_Capabilities_for_Embodied_Task_Planning_CVPR_2026_paper.pdf" target="_blank">https://openaccess.thecvf.com/content/CVPR2026/papers/Xu_RoboAgent_Chaining_Basic_Capabilities_for_Embodied_Task_Planning_CVPR_2026_paper.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[48] Bi et al., Motus. CVPR 2026. <a href="https://openaccess.thecvf.com/content/CVPR2026/papers/Bi_Motus_A_Unified_Latent_Action_World_Model_CVPR_2026_paper.pdf" target="_blank">https://openaccess.thecvf.com/content/CVPR2026/papers/Bi_Motus_A_Unified_Latent_Action_World_Model_CVPR_2026_paper.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[49] PaperNotes, CVPR 2026 Robotics &amp; Embodied AI. <a href="https://en.papernotes.org/CVPR2026/robotics/" target="_blank">https://en.papernotes.org/CVPR2026/robotics/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[50] HERA. arXiv:2604.00901. <a href="https://arxiv.org/pdf/2604.00901v2" target="_blank">https://arxiv.org/pdf/2604.00901v2</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[51] Forasoft, LangGraph vs CrewAI vs AutoGen. <a href="https://www.forasoft.com/learn/ai-for-video-engineering/articles-ai/langgraph-vs-crewai-vs-autogen-agent-frameworks" target="_blank">https://www.forasoft.com/learn/ai-for-video-engineering/articles-ai/langgraph-vs-crewai-vs-autogen-agent-frameworks</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[52] Andrew.ooo, LangGraph 1.2 vs CrewAI 1.14. <a href="https://andrew.ooo/answers/langgraph-1-2-vs-crewai-1-14-vs-mastra-may-2026/" target="_blank">https://andrew.ooo/answers/langgraph-1-2-vs-crewai-1-14-vs-mastra-may-2026/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[53] AliceLabs, Best AI Agent Frameworks 2026. <a href="https://alicelabs.ai/en/insights/best-ai-agent-frameworks-2026" target="_blank">https://alicelabs.ai/en/insights/best-ai-agent-frameworks-2026</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[54] AIWiki, Computer-Use Agent. <a href="https://aiwiki.ai/wiki/computer-use_agent" target="_blank">https://aiwiki.ai/wiki/computer-use_agent</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[55] DataCamp, Muse Spark 1.1. <a href="https://www.datacamp.com/id/blog/muse-spark-1-1" target="_blank">https://www.datacamp.com/id/blog/muse-spark-1-1</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[56] FelloAI, Claude Cowork Guide. <a href="https://felloai.com/ko/claude-cowork-guide/" target="_blank">https://felloai.com/ko/claude-cowork-guide/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[57] Zylos AI, Graph-Based Agent Workflow Orchestration. <a href="https://zylos.ai/en/research/2026-04-14-graph-based-agent-workflow-orchestration-production/" target="_blank">https://zylos.ai/en/research/2026-04-14-graph-based-agent-workflow-orchestration-production/</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[58] Emerging AI Engineering. <a href="https://media.aivoid.dev/pdfs/emerging-ai-engineering-agents-orchestration-and-ai-native-systems_20260504181824.pdf" target="_blank">https://media.aivoid.dev/pdfs/emerging-ai-engineering-agents-orchestration-and-ai-native-systems_20260504181824.pdf</a></p><p style="font-family: -apple-system, BlinkMacSystemFont, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Helvetica Neue&#34;, sans-serif;letter-spacing: normal;text-align: start;background-color: rgb(250, 250, 247);font-size: 13px;line-height: 1.8;color: rgb(107, 114, 128);margin-bottom: 4px;">[59] Hugging Face, Smolagents. <a href="https://smolagents.org/hi/" target="_blank">https://smolagents.org/hi/</a></p><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=6fb16a6a&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486132%26idx%3D1%26sn%3Df769d3ac7cf5df69da11227e88301e21">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Tue, 21 Jul 2026 08:01:00 +0800</pubDate>
    </item>
    <item>
      <title>2026年网安/软工/AI顶会基于Agent 的漏洞挖掘、验证、利用与修复的论文汇编</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486128&amp;idx=1&amp;sn=791ba32939688d0a2e21fdf74b02af4b</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-07-20 11:02</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=74bb4ae6&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sxHncHlPk5DKwSDc3shRE3tNxVs7xrZHibib7bbBgvy1pKUUfC3hCvibzy5E50e13ia9tvH2pgwzV1oTSO02x2nXIZgcUAtrAADNrw%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <header class="cover" style="padding: 5rem 1.5rem 4rem;background: linear-gradient(135deg, rgb(26, 35, 50) 0%, rgb(45, 27, 46) 50%, rgb(30, 58, 95) 100%);color: rgb(255, 255, 255);overflow: hidden;font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="cover-inner" style="margin-right: auto;margin-left: auto;max-width: 980px;z-index: 1;"><div class="eyebrow" style="margin-bottom: 1.2rem;font-family: JetBrainsMono, monospace;font-size: 0.82rem;letter-spacing: 0.18em;text-transform: uppercase;color: rgb(252, 165, 165);">LITERATURE SURVEY · 2026</div><h1 style="margin-bottom: 1rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.9rem;font-weight: 700;line-height: 1.25;letter-spacing: -0.01em;">2026 年顶会 Agent 漏洞研究论文汇编</h1><p class="subtitle" style="font-size: 1.15rem;color: rgb(203, 213, 225);max-width: 680px;line-height: 1.6;">系统收集 2026 年网安四大顶会、软工顶会、安全测试顶会与 AI 顶会中,所有关于智能体(Agent)用于漏洞挖掘、验证、利用与修复的论文,涵盖标题、作者、机构、摘要与下载链接。</p><div class="meta" style="margin-top: 2rem;display: flex;flex-wrap: wrap;gap: 1.5rem 2.5rem;font-size: 0.9rem;color: rgb(148, 163, 184);font-family: JetBrainsMono, monospace;"><div><span style="color: rgb(255, 255, 255);font-weight: 700;">检索日期</span> 2026-07-19</div><div><span style="color: rgb(255, 255, 255);font-weight: 700;">覆盖会议</span> 11 个顶会</div><div><span style="color: rgb(255, 255, 255);font-weight: 700;">收录论文</span> 26 篇正文 + 8 篇预印本参考</div></div></div></header><div class="stats-strip" style="padding: 2rem 1.5rem;background: var(--bg2);border-bottom: 1px solid var(--rule);color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="stats-grid" style="margin-right: auto;margin-left: auto;max-width: 980px;display: grid;grid-template-columns: repeat(2, 1fr);gap: 1rem;"><div class="stat-card" style="text-align: center;"><div class="stat-num" style="font-family: InstrumentSans, sans-serif;font-size: 2.4rem;font-weight: 700;color: var(--accent);line-height: 1;">12</div><div class="stat-label" style="margin-top: 0.4rem;font-size: 0.82rem;color: var(--muted);letter-spacing: 0.02em;">网安四大顶会论文</div></div><div class="stat-card" style="text-align: center;"><div class="stat-num" style="font-family: InstrumentSans, sans-serif;font-size: 2.4rem;font-weight: 700;color: var(--accent);line-height: 1;">4</div><div class="stat-label" style="margin-top: 0.4rem;font-size: 0.82rem;color: var(--muted);letter-spacing: 0.02em;">软工 / 安全测试顶会</div></div><div class="stat-card" style="text-align: center;"><div class="stat-num" style="font-family: InstrumentSans, sans-serif;font-size: 2.4rem;font-weight: 700;color: var(--accent);line-height: 1;">10</div><div class="stat-label" style="margin-top: 0.4rem;font-size: 0.82rem;color: var(--muted);letter-spacing: 0.02em;">AI 顶会(含研讨会)</div></div><div class="stat-card" style="text-align: center;"><div class="stat-num" style="font-family: InstrumentSans, sans-serif;font-size: 2.4rem;font-weight: 700;color: var(--accent);line-height: 1;">8</div><div class="stat-label" style="margin-top: 0.4rem;font-size: 0.82rem;color: var(--muted);letter-spacing: 0.02em;">相关预印本参考</div></div></div></div><p style="padding-top: 2.5rem;padding-bottom: 2.5rem;color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="container" style="margin-right: auto;margin-left: auto;padding-right: 1.5rem;padding-left: 1.5rem;max-width: 980px;"><div class="section-head" style="margin-bottom: 2.5rem;"><div class="kicker" style="margin-bottom: 0.6rem;font-family: JetBrainsMono, monospace;font-size: 0.78rem;letter-spacing: 0.15em;text-transform: uppercase;color: var(--accent);">OVERVIEW · 概述</div><h2 style="margin-bottom: 0.6rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.5rem;font-weight: 700;color: var(--ink);line-height: 1.3;">研究背景与检索范围</h2><p class="lead" style="color: var(--muted);font-size: 1rem;max-width: 720px;">大语言模型(LLM)智能体正在重塑漏洞研究的全生命周期。本报告系统检索 2026 年网络安全、软件工程与人工智能领域的顶级会议,汇总其中&#34;以 Agent 为执行主体、用于漏洞挖掘/验证/利用/修复&#34;的论文。</p></div><p>本次检索覆盖 <strong>11 个目标会议</strong>:网安四大顶会(IEEE S&amp;P、ACM CCS、USENIX Security、NDSS)、软工顶会(ICSE、FSE、ASE)、安全测试顶会(ISSTA)、AI 顶会(NeurIPS、ICML、ICLR、AAAI、IJCAI、CVPR)。检索时点为 2026 年 7 月 19 日,各会议状态如下:</p><div class="table-wrap" style="margin-top: 1.5rem;margin-bottom: 1.5rem;overflow: auto;max-height: 600px;border: 1px solid var(--rule);border-radius: 8px;"><table style="margin-bottom: 0px;width: 720px;font-size: 0.85rem;min-width: 720px;"><thead style="background: var(--ink);color: rgb(255, 255, 255);top: 0px;"><tr><th style="padding: 0.7rem 0.8rem;text-align: left;font-family: InstrumentSans, sans-serif;font-size: 0.82rem;letter-spacing: 0.02em;">会议</th><th style="padding: 0.7rem 0.8rem;text-align: left;font-family: InstrumentSans, sans-serif;font-size: 0.82rem;letter-spacing: 0.02em;">举办时间</th><th style="padding: 0.7rem 0.8rem;text-align: left;font-family: InstrumentSans, sans-serif;font-size: 0.82rem;letter-spacing: 0.02em;">论文公布情况</th><th style="padding: 0.7rem 0.8rem;text-align: left;font-family: InstrumentSans, sans-serif;font-size: 0.82rem;letter-spacing: 0.02em;">相关论文数</th></tr></thead><tbody><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">NDSS 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 2 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布(完整)</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">4 篇</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">IEEE S&amp;P 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 5 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布(完整)</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">3 篇</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">USENIX Security 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 8 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">Cycle 1 已公布,Cycle 2 待公开</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2 篇</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ACM CCS 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 11 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">Cycle 1 通知(7/17),完整列表待公开</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">3 篇(+1 待确认)</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ICSE 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 4 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布(完整)</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2 篇</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">FSE 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 7 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2 篇</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ASE 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 10 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">通知刚发,proceedings 未公开</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">待核实</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ISSTA 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 10 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">通知已发,proceedings 未公开</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">待核实</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ICLR 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 5 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布(完整)</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">3 篇主会 + 2 篇研讨会</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">ICML 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 7 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">1 篇主会 + 2 篇研讨会</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">AAAI 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 1 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">1 篇研讨会</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">CVPR 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 6 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">已公布</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">0 篇(无相关主题)</td></tr><tr><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">IJCAI 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 8 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">即将召开,列表未完全公开</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">待核实</td></tr><tr style="background: var(--bg);"><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">NeurIPS 2026</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">2026 年 12 月</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">尚未召开(通知 9/24)</td><td style="padding: 0.6rem 0.8rem;border-bottom: 1px solid var(--rule);vertical-align: top;">无法检索</td></tr></tbody></table></div><figure class="chart-figure" style="margin-top: 2rem;margin-bottom: 2rem;padding: 1.5rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;"><figcaption style="margin-bottom: 1rem;font-size: 0.95rem;font-weight: 700;color: var(--ink);">各会议相关论文数量分布</figcaption><div echarts_instance_="ec_1784510562477" style="width: 362.545px;min-height: 380px;"><div style="overflow: hidden;width: 363px;height: 380px;cursor: default;"><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" version="1.1" baseProfile="full" width="363" height="380" style="left: 0px;top: 0px;user-select: none;"><rect width="363" height="380" x="0" y="0" fill="none"></rect><g><path d="M25.16 338.5L348.48 338.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><path d="M25.16 280.5L348.48 280.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><path d="M25.16 221.5L348.48 221.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><path d="M25.16 163.5L348.48 163.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><path d="M25.16 104.5L348.48 104.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><path d="M25.16 45.5L348.48 45.5" fill="transparent" stroke="#e3e6ec" stroke-dasharray="4,2"></path><text dominant-baseline="central" text-anchor="middle" y="-5.499961853027344" transform="translate(25.16 30.6)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">论文数量</text><path d="M25.16 338.5L348.48 338.5" fill="transparent" stroke="#e3e6ec" stroke-linecap="round"></path><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 338.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">0</text><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 280.0001)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">1</text><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 221.4001)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">2</text><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 162.8001)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">3</text><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 104.2)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">4</text><text dominant-baseline="central" text-anchor="end" transform="translate(17.16 45.6)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">5</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(43.1222 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">NDSS</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(79.0466 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">S&amp;P</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(114.9711 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">USENIX</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(150.8955 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">CCS</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(186.82 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">ICSE</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(222.7444 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">FSE</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(258.6689 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">ICLR</text><text dominant-baseline="central" text-anchor="middle" y="16.49988555908203" transform="translate(258.6689 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">(含研讨会)</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(294.5933 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">ICML</text><text dominant-baseline="central" text-anchor="middle" y="16.49988555908203" transform="translate(294.5933 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">(含研讨会)</text><text dominant-baseline="central" text-anchor="middle" y="5.499961853027344" transform="translate(330.5178 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">AAAI</text><text dominant-baseline="central" text-anchor="middle" y="16.49988555908203" transform="translate(330.5178 346.6002)" fill="#5a6678" style="font-size: 11px;font-family: sans-serif;">(研讨会)</text><path d="M37.7818 104.2L48.4625 104.2A4 4 0 0 1 52.4625 108.2L52.4625 338.6002L33.7818 338.6002L33.7818 108.2A4 4 0 0 1 37.7818 104.2" fill="#b91c1c"></path><path d="M73.7063 162.8001L84.387 162.8001A4 4 0 0 1 88.387 166.8001L88.387 338.6002L69.7063 338.6002L69.7063 166.8001A4 4 0 0 1 73.7063 162.8001" fill="#b91c1c"></path><path d="M109.6307 221.4001L120.3114 221.4001A4 4 0 0 1 124.3114 225.4001L124.3114 338.6002L105.6307 338.6002L105.6307 225.4001A4 4 0 0 1 109.6307 221.4001" fill="#b91c1c"></path><path d="M145.5552 162.8001L156.2359 162.8001A4 4 0 0 1 160.2359 166.8001L160.2359 338.6002L141.5552 338.6002L141.5552 166.8001A4 4 0 0 1 145.5552 162.8001" fill="#b91c1c"></path><path d="M181.4796 221.4001L192.1603 221.4001A4 4 0 0 1 196.1603 225.4001L196.1603 338.6002L177.4796 338.6002L177.4796 225.4001A4 4 0 0 1 181.4796 221.4001" fill="#1e40af"></path><path d="M217.4041 221.4001L228.0848 221.4001A4 4 0 0 1 232.0848 225.4001L232.0848 338.6002L213.4041 338.6002L213.4041 225.4001A4 4 0 0 1 217.4041 221.4001" fill="#1e40af"></path><path d="M253.3285 45.6L264.0092 45.6A4 4 0 0 1 268.0092 49.6L268.0092 338.6002L249.3285 338.6002L249.3285 49.6A4 4 0 0 1 253.3285 45.6" fill="#1e40af"></path><path d="M289.253 162.8001L299.9337 162.8001A4 4 0 0 1 303.9337 166.8001L303.9337 338.6002L285.253 338.6002L285.253 166.8001A4 4 0 0 1 289.253 162.8001" fill="#1e40af"></path><path d="M325.1774 280.0001L335.8581 280.0001A4 4 0 0 1 339.8581 284.0001L339.8581 338.6002L321.1774 338.6002L321.1774 284.0001A4 4 0 0 1 325.1774 280.0001" fill="#1e40af"></path><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(43.1222 99.2)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">4</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(79.0466 157.8001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">3</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(114.9711 216.4001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">2</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(150.8955 157.8001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">3</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(186.82 216.4001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">2</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(222.7444 216.4001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">2</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(258.6689 40.6)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">5</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(294.5933 157.8001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">3</text><text dominant-baseline="central" text-anchor="middle" y="-6.0000457763671875" transform="translate(330.5178 275.0001)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 700;">1</text></g></svg>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div></div></div><div class="chart-sub" style="margin-top: 0.6rem;font-size: 0.82rem;color: var(--muted);">注:ICLR 与 ICML 含研讨会论文;CCS 含 1 篇归属待确认论文。</div></figure><figure class="chart-figure" style="margin-top: 2rem;margin-bottom: 2rem;padding: 1.5rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;"><figcaption style="margin-bottom: 1rem;font-size: 0.95rem;font-weight: 700;color: var(--ink);">论文按研究任务类型分布(一篇论文可属于多个类型)</figcaption><div echarts_instance_="ec_1784510562478" style="width: 362.545px;min-height: 340px;"><div style="overflow: hidden;width: 363px;height: 340px;"><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" version="1.1" baseProfile="full" width="363" height="340" style="left: 0px;top: 0px;user-select: none;"><rect width="363" height="340" x="0" y="0" fill="none"></rect><g><path d="M286.7622 125.4808L298.372 122.4456L306.372 122.4456" fill="transparent" stroke="#e3e6ec"></path><path d="M165.6586 260.6406L163.9114 272.5127L155.9114 272.5127" fill="transparent" stroke="#e3e6ec"></path><path d="M74.5098 172.7621L62.7094 174.9417L54.7094 174.9417" fill="transparent" stroke="#e3e6ec"></path><path d="M121.4677 62.261L114.8465 52.2531L106.8465 52.2531" fill="transparent" stroke="#e3e6ec"></path><path d="M181.5 44.2A108.8 108.8 0 0 1 234.7487 247.879L213.1164 209.3344A64.6 64.6 0 0 0 181.5 88.4Z" fill="#b91c1c" stroke="#ffffff" stroke-width="2" stroke-linejoin="round"></path><path d="M234.7487 247.879A108.8 108.8 0 0 1 103.1745 228.5153L134.9942 197.8372A64.6 64.6 0 0 0 213.1164 209.3344Z" fill="#1e40af" stroke="#ffffff" stroke-width="2" stroke-linejoin="round"></path><path d="M103.1745 228.5153A108.8 108.8 0 0 1 81.3663 110.4478L122.0456 127.7346A64.6 64.6 0 0 0 134.9942 197.8372Z" fill="#475569" stroke="#ffffff" stroke-width="2" stroke-linejoin="round"></path><path d="M81.3663 110.4478A108.8 108.8 0 0 1 181.5 44.2L181.5 88.4A64.6 64.6 0 0 0 122.0456 127.7346Z" fill="#15803d" stroke="#ffffff" stroke-width="2" stroke-linejoin="round"></path><text dominant-baseline="central" text-anchor="start" y="-6.0000457763671875" transform="translate(311.372 122.4456)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">漏洞...</text><text dominant-baseline="central" text-anchor="start" xml:space="preserve" y="6.0000457763671875" transform="translate(311.372 122.4456)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">18 篇</text><text dominant-baseline="central" text-anchor="end" y="-6.0000457763671875" transform="translate(150.9114 272.5127)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">漏洞验证</text><text dominant-baseline="central" text-anchor="end" xml:space="preserve" y="6.0000457763671875" transform="translate(150.9114 272.5127)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">9 篇</text><text dominant-baseline="central" text-anchor="end" y="-6.0000457763671875" transform="translate(49.7094 174.9417)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">漏洞...</text><text dominant-baseline="central" text-anchor="end" xml:space="preserve" y="6.0000457763671875" transform="translate(49.7094 174.9417)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">8 篇</text><text dominant-baseline="central" text-anchor="end" y="-6.0000457763671875" transform="translate(101.8465 52.2531)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">漏洞修复</text><text dominant-baseline="central" text-anchor="end" xml:space="preserve" y="6.0000457763671875" transform="translate(101.8465 52.2531)" fill="#1a2332" style="font-size: 12px;font-family: sans-serif;font-weight: 600;">8 篇</text><path d="M-5 -5l334.0015 0l0 24l-334.0015 0Z" transform="translate(19.4993 311)" fill="rgb(0,0,0)" fill-opacity="0" stroke="#ccc" stroke-width="0"></path><path d="M3 0L9 0A3 3 0 0 1 12 3L12 9A3 3 0 0 1 9 12L3 12A3 3 0 0 1 0 9L0 3A3 3 0 0 1 3 0" transform="translate(20.4993 312)" fill="#b91c1c" stroke="#ffffff" stroke-width="2" stroke-linecap="butt" stroke-miterlimit="10" stroke-linejoin="round"></path><text dominant-baseline="central" text-anchor="start" x="17" y="6" transform="translate(20.4993 312)" fill="#5a6678" style="font-size: 12px;font-family: sans-serif;">漏洞挖掘</text><path d="M-1 -1l66.0004 0l0 14l-66.0004 0Z" transform="translate(20.4993 312)" fill="transparent"></path><path d="M3 0L9 0A3 3 0 0 1 12 3L12 9A3 3 0 0 1 9 12L3 12A3 3 0 0 1 0 9L0 3A3 3 0 0 1 3 0" transform="translate(106.4996 312)" fill="#1e40af" stroke="#ffffff" stroke-width="2" stroke-linecap="butt" stroke-miterlimit="10" stroke-linejoin="round"></path><text dominant-baseline="central" text-anchor="start" x="17" y="6" transform="translate(106.4996 312)" fill="#5a6678" style="font-size: 12px;font-family: sans-serif;">漏洞验证</text><path d="M-1 -1l66.0004 0l0 14l-66.0004 0Z" transform="translate(106.4996 312)" fill="transparent"></path><path d="M3 0L9 0A3 3 0 0 1 12 3L12 9A3 3 0 0 1 9 12L3 12A3 3 0 0 1 0 9L0 3A3 3 0 0 1 3 0" transform="translate(192.5 312)" fill="#475569" stroke="#ffffff" stroke-width="2" stroke-linecap="butt" stroke-miterlimit="10" stroke-linejoin="round"></path><text dominant-baseline="central" text-anchor="start" x="17" y="6" transform="translate(192.5 312)" fill="#5a6678" style="font-size: 12px;font-family: sans-serif;">漏洞利用</text><path d="M-1 -1l66.0004 0l0 14l-66.0004 0Z" transform="translate(192.5 312)" fill="transparent"></path><path d="M3 0L9 0A3 3 0 0 1 12 3L12 9A3 3 0 0 1 9 12L3 12A3 3 0 0 1 0 9L0 3A3 3 0 0 1 3 0" transform="translate(278.5004 312)" fill="#15803d" stroke="#ffffff" stroke-width="2" stroke-linecap="butt" stroke-miterlimit="10" stroke-linejoin="round"></path><text dominant-baseline="central" text-anchor="start" x="17" y="6" transform="translate(278.5004 312)" fill="#5a6678" style="font-size: 12px;font-family: sans-serif;">漏洞修复</text><path d="M-1 -1l66.0004 0l0 14l-66.0004 0Z" transform="translate(278.5004 312)" fill="transparent"></path></g></svg>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div>&lt;div style=&#34;position: absolute !important; visibility: hidden !important; border-width: 0px !important; user-select: none !important; width: 0px !important; height: 0px !important; <div absolute="" visibility:="" hidden="" border-width:="" px="" user-select:="" none="" width:="" height:="" inset:="" auto="">&#34;&gt;</div></div></div><div class="chart-sub" style="margin-top: 0.6rem;font-size: 0.82rem;color: var(--muted);">任务类型说明:挖掘(漏洞发现/检测)、验证(漏洞确认/PoV/PoC 生成)、利用(漏洞利用/exploit 生成/渗透测试)、修复(漏洞补丁/程序修复)。</div></figure><div class="notice notice-warn" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: rgba(161, 98, 7, 0.08);border-left-color: rgb(161, 98, 7);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: rgb(161, 98, 7);">重要说明:</strong> 所有论文信息均来自会议官网、ACM Digital Library、OpenReview、arXiv 等公开来源的真实检索结果,未编造任何论文。部分 2026 年会议(ASE、ISSTA、IJCAI、NeurIPS)因尚未召开或 proceedings 未公开,可能存在遗漏,建议在论文列表正式公开后复查。研讨会(Workshop)论文已明确标注,与主会论文区分。</div></div></p><p style="padding-top: 2.5rem;padding-bottom: 2.5rem;border-top: 1px solid var(--rule);color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="container" style="margin-right: auto;margin-left: auto;padding-right: 1.5rem;padding-left: 1.5rem;max-width: 980px;"><div class="section-head" style="margin-bottom: 2.5rem;"><div class="kicker" style="margin-bottom: 0.6rem;font-family: JetBrainsMono, monospace;font-size: 0.78rem;letter-spacing: 0.15em;text-transform: uppercase;color: var(--accent);">PART 1 · 网安四大顶会</div><h2 style="margin-bottom: 0.6rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.5rem;font-weight: 700;color: var(--ink);line-height: 1.3;">网络安全四大顶会(NDSS / S&amp;P / USENIX / CCS)</h2><p class="lead" style="color: var(--muted);font-size: 1rem;max-width: 720px;">网安四大顶会是漏洞研究最核心的发表阵地。2026 年四大会议共检索到 12 篇相关论文(NDSS 4 篇、S&amp;P 3 篇、USENIX 2 篇、CCS 3 篇),另有 1 篇 CCS 归属待确认。</p></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">NDSS 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">网络与分布式系统安全研讨会</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 2 月 · 已完整公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">01</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">FirmAgent: Leveraging Fuzzing to Assist LLM Agents with IoT Firmware Vulnerability Discovery</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">FirmAgent:利用模糊测试辅助 LLM 智能体发现 IoT 固件漏洞</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">NDSS 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Jiangan Ji, Chao Zhang, Shuitao Gan, Lin Jian, Hangtian Liu, Tieming Liu, Lei Zheng, Zhipeng Jia</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（信息工程大学）</span> · <span class="affil" style="color: var(--accent2);">（清华大学网络科学与网络空间研究院/JCSS）</span> · <span class="affil" style="color: var(--accent2);">（先进计算与智能工程实验室）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>FirmAgent 是首个混合解决方案,利用模糊测试(fuzzing)辅助 LLM 智能体发现 IoT 固件中的漏洞。其设计动机来自一个关键观察:模糊测试能准确识别固件中与输入相关的代码点,而静态分析能从这些代码点出发深入分析程序路径。FirmAgent 利用模糊测试收集运行时输入点(即污点源)并重建潜在漏洞路径,然后应用一个 LLM 智能体(污点传播智能体)沿潜在路径执行上下文感知的污点分析,再由另一个 LLM 智能体(PoC 生成智能体)将模糊测试生成的测试用例精炼为可验证漏洞的 PoC。在 14 个真实 IoT 固件上的评估中,FirmAgent 以 91% 的精度识别出 182 个漏洞,包含 140 个此前未知漏洞,其中 17 个已获 CVE 编号。它在检测能力和精度上均显著超越 SOTA 工具。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1943-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1943-paper.pdf</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">代码</span><a href="https://github.com/vul337/FirmAgent.git" target="_blank">https://github.com/vul337/FirmAgent.git</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">02</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Context Relay for Long-Running Penetration-Testing Agents (CHAP)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">面向长期运行渗透测试智能体的上下文中继</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">NDSS 2026</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Marius Vangeli, Joel Brynielsson, Mika Cohen, Farzad Kamrani</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（瑞典皇家工学院 KTH Royal Institute of Technology）</span> · <span class="affil" style="color: var(--accent2);">（瑞典国防研究院 FOI Swedish Defence Research Agency）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>虽然基于 LLM 的渗透测试正在快速进步,但自主智能体仍难以完成较长时间的多阶段漏洞利用。随着智能体执行侦察、尝试利用漏洞、在系统间横向移动,token 上下文窗口会被探索过程和失败尝试填满,导致决策质量下降。本文提出 CHAP(Context Handoff for Autonomous Penetration Testing),一种为 LLM 驱动智能体设计的上下文中继系统。CHAP 使智能体能够通过将积累的知识作为紧凑协议传递给新的智能体实例来维持长期运行的渗透测试。在 AutoPenBench 扩展版基准上的评估显示,CHAP 针对 11 个真实漏洞,将每次运行的成功率从 27.3% 提升至 36.4%,同时相比基线智能体减少 32.4% 的 token 开销。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://www.ndss-symposium.org/wp-content/uploads/lastx2026-42.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/lastx2026-42.pdf</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">03</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">AWE: Adaptive Agents for Dynamic Web Penetration Testing</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">AWE:面向动态 Web 渗透测试的自适应智能体</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">NDSS 2026</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Akshat Singh Jaswal, Ashish Baghel</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（Stux Labs）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>现代 Web 应用越来越多地通过 AI 辅助开发和快速 no-code 部署流水线产生,加剧了软件速度与安全工具适应性之间的差距。模式驱动的扫描器无法理解新场景,而新兴的基于 LLM 的渗透测试工具依赖无约束的探索,导致成本高、行为不稳定、可复现性差。本文提出 AWE,一种记忆增强的多智能体框架,用于自主 Web 渗透测试,将结构化、特定漏洞类型的分析流水线嵌入到轻量级 LLM 编排层中。与通用智能体不同,AWE 将上下文感知的 payload 变异和生成与持久记忆及浏览器验证相结合,产生确定性的、漏洞利用驱动的结果。在 104 道题的 XBOW 基准上评估,AWE 在注入类漏洞上取得显著成果——XSS 成功率 87%(较 MAPTA 提升 30.5%)、盲注 SQLi 成功率 66.7%(提升 33.3%)——同时比 MAPTA 更快、更便宜、更高效。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://www.ndss-symposium.org/wp-content/uploads/lastx2026-37.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/lastx2026-37.pdf</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">代码</span><a href="https://github.com/stuxlabs/AWE" target="_blank">https://github.com/stuxlabs/AWE</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">04</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">BSFuzzer: Context-Aware Semantic Fuzzing for BLE Logic Flaw Detection</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">BSFuzzer:面向 BLE 逻辑缺陷检测的上下文感知语义模糊测试</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">NDSS 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Ting Yang, Yue Qin, Lan Zhang, Zhiyuan Fu, Junfan Chen, Jice Wang, Shangru Zhao, Qi Li, Ruidong Li, He Wang, Yuqing Zhang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（西安电子科技大学）</span> · <span class="affil" style="color: var(--accent2);">（金泽大学）</span> · <span class="affil" style="color: var(--accent2);">（中央财经大学）</span> · <span class="affil" style="color: var(--accent2);">（北亚利桑那大学）</span> · <span class="affil" style="color: var(--accent2);">（海南大学）</span> · <span class="affil" style="color: var(--accent2);">（中国科学院大学）</span> · <span class="affil" style="color: var(--accent2);">（清华大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>蓝牙低功耗(BLE)是现代互联设备的基础通信标准,但其复杂设计引入了细微的逻辑缺陷(如字段误解或无效状态转换),可导致认证绕过、未授权控制或 DoS 攻击,且常规避传统模糊测试和形式化分析。本文提出 BSFuzzer,一种由 Bluetooth Core Specification 引导的黑盒、上下文感知语义模糊测试框架。BSFuzzer 使用 LLM 智能体对蓝牙规范进行语义解析,从文本、图表和上下文中提取状态机和包语义;然后生成两类变异:违反协议规则的字段级变异和破坏关键状态转换的状态级变异。LLM 智能体还用于根据预期行为验证响应,能检测传统 fuzzer 无法发现的细微逻辑缺陷。在 19 款真实 BLE 设备(9 个 SoC 模块、10 款智能手机)上的评估中,BSFuzzer 发现 36 个安全问题,含 34 个此前未知 bug,其中 9 个获 CVE 编号。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f94-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f94-paper.pdf</a></div></div></article></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">S&amp;P 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">IEEE 安全与隐私研讨会(Oakland)</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 5 月 · 已完整公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">05</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Agentic Concolic Execution</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">智能体驱动的 Concolic 执行</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">S&amp;P 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Zhengxiong Luo, Huan Zhao, Dylan Wolff, Cristian Cadar, Abhik Roychoudhury</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（新加坡国立大学 National University of Singapore）</span> · <span class="affil" style="color: var(--accent2);">（英国帝国理工学院 Imperial College London）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>该论文提出&#34;Agentic Concolic Execution&#34;,将 LLM 智能体与 concolic 执行(混合具体执行与符号执行)相结合,用于程序分析与漏洞发现。其核心思想是让 LLM 智能体动态指导 concolic 执行引擎,利用智能体的语义推理能力突破传统符号执行面临的约束求解瓶颈,从而能够探索更深的程序路径并触发更多漏洞。这是将 LLM agent 范式引入符号执行/测试生成领域的工作,展示了智能体在自动化漏洞发现中的潜力。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://abhikrc.com/pdf/SP26.pdf" target="_blank">https://abhikrc.com/pdf/SP26.pdf</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">项目</span><a href="https://concollmic.github.io/" target="_blank">https://concollmic.github.io/</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">06</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">GONDAR: Contextualizing Sink Knowledge for Java Vulnerability Discovery</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">GONDAR:上下文化 Sink 知识用于 Java 漏洞发现</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">S&amp;P 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Fabian Fleischer, Cen Zhang, Joonun Jang, Jeongin Cho, Meng Xu, Taesoo Kim</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国佐治亚理工学院 Georgia Institute of Technology）</span> · <span class="affil" style="color: var(--accent2);">（三星研究院 Samsung Research）</span> · <span class="affil" style="color: var(--accent2);">（加拿大滑铁卢大学 University of Waterloo）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>Java 应用易受安全敏感 API(如文件操作导致路径遍历、反序列化导致 RCE)的不安全使用所引发的漏洞。现有 fuzzer 在很大程度上忽视了这种特定漏洞知识,限制了其有效性。本文提出 GONDAR,一种以 sink 为中心的 fuzzing 框架,系统性利用 sink API 语义进行有针对性的漏洞发现。GONDAR 首先通过 CWE 特定扫描结合 LLM 辅助的静态过滤识别可达且可利用的 sink 调用点;然后部署两个专门智能体与覆盖率引导的 fuzzer 协作:(1)探索智能体(exploration agent)通过迭代求解路径约束生成到达目标调用点的输入;(2)利用智能体(exploitation agent)通过推理和满足漏洞触发条件来合成 PoC exploit。智能体与 fuzzer 持续交换种子和运行时反馈,互相补充。在真实 Java 基准上,GONDAR 发现的漏洞数量是 SOTA Java fuzzer Jazzer 的 4 倍。GONDAR 还在 DARPA AI 网络挑战赛(AIxCC)中表现优异,并已集成到 Linux 基金会 OpenSSF 的 OSS-CRS 沙箱项目中。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/pdf/2604.01645v2" target="_blank">https://arxiv.org/pdf/2604.01645v2</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">07</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Neo: Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Neo:通过智能体化程序分析检测多语言微服务中的权限提升漏洞</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">S&amp;P 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Penghui Li, Hong Yau Chong, Yinzhi Cao, Junfeng Yang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国约翰斯·霍普金斯大学 Johns Hopkins University）</span> · <span class="affil" style="color: var(--accent2);">（美国哥伦比亚大学 Columbia University）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>微服务架构在云系统中广泛采用,但其在权限和权限控制方面引入了显著复杂性,存在权限提升风险——攻击者可获得未授权的资源或操作访问。检测此类漏洞极具挑战,因为存在复杂的跨服务交互、多语言代码库以及多样化的特权操作和权限检查。本文提出 Neo,一个结合 LLM 与经典程序分析的智能体化程序分析框架。Neo 利用基于 LLM 的智能体动态生成分析计划、调整代码搜索策略并验证语义。我们开发了代码搜索原语,使 Neo 能在跨服务和跨语言间进行可扩展、灵活的代码探索。在 25 个开源微服务应用(7 种编程语言、620 万行代码)上的评估中,Neo 发现 24 个零日权限提升漏洞,在 ground-truth 数据集上达到 81.0% 精度和 85.0% 召回率,相比现有程序分析和 agentic 解决方案均有显著改进。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/html/2605.15569v1" target="_blank">https://arxiv.org/html/2605.15569v1</a></div></div></article></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">USENIX 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">USENIX 安全研讨会</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 8 月 · Cycle 1 已公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">08</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">iFinder: Understanding Implicit Trust Errors in Core Carrier Networks through Multi-Agent Flaw Discovery and Analysis</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">iFinder:通过多智能体缺陷发现与分析理解核心承载网中的隐式信任错误</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">USENIX 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Ziyu Lin, Ziting Wang, Xinfeng Li, Wei Dong, XiaoFeng Wang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（新加坡南洋理工大学 Nanyang Technological University）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>蜂窝核心网(CN)是关键基础设施,但其内部安全模型长期依赖物理隔离——核心组件之间的接口通常在假定信任域内运行。随着 CN 向云原生部署转型,此假设弱化,使外部对手能触及此前内部的接口。通过对开源 CN 实现中 GitHub issues 报告的安全缺陷进行根因分析,我们发现了 CN 组件间反复出现的&#34;盲目信任&#34;模式——组件可能省略语法验证、未强制执行语义不变式,或在未检查可用性的情况下分配资源。一旦内部接口变得可达,这些弱点可导致拒绝服务、会话劫持等严重后果。我们称之为&#34;隐式信任错误&#34;(iTrue)。为检测 iTrue 并理解其安全影响,我们设计了 iFinder——一个 LLM 驱动的多智能体系统,可总结已知缺陷、将其提炼为检测模式,并应用这些模式发现 CN 实现中的新 iTrue。为抑制 LLM 幻觉,我们构建了一种交叉检查 3GPP 规范和 CN 代码的创新策略;还开发了利用 LLM 生成 PoC exploit 并通过自动执行迭代精炼的技术。在 7 个主流开源 CN 实现上运行 iFinder,发现 84 个此前未知漏洞,其中 83 个已被确认,81 个已分配 CVE 编号。一个会话劫持缺陷已在真实商业 5G 核心网络上得到确认。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/pdf/2607.10315" target="_blank">https://arxiv.org/pdf/2607.10315</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">09</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">SoK: DARPA&#39;s AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">SoK:DARPA AI 网络挑战赛(AIxCC)——竞赛设计、架构与经验教训</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">USENIX 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">综述</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">综述</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Michael Pelican, David J. Musliner, Jeff Huang, Jon Silliman, Mikel Mcdaniel, Jefferson Casavant, Isaac Goldthwaite, Nicholas Vidovich, Matthew Lehman, Taesoo Kim</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国佐治亚理工学院 Georgia Institute of Technology）</span> · <span class="affil" style="color: var(--accent2);">（及其他 AIxCC 参赛团队机构）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>DARPA AI 网络挑战赛(AIxCC,2023-2025)是迄今为止规模最大的竞赛,旨在构建完全自主的网络推理系统(CRS),利用 AI(尤其是 LLM)在真实开源软件中发现并修复漏洞。本文对 AIxCC 进行首次系统性分析。基于设计文档、源代码、执行追踪以及与组织者和参赛团队的讨论,论文审视了竞赛结构和关键设计决策,刻画了决赛 CRS 的架构方法,并分析了超越最终成绩单的竞赛结果。分析揭示了真正驱动 CRS 性能的因素、团队所取得的真正技术进步,以及仍待解决的限制。论文最后总结了组织未来竞赛的经验教训,以及部署自主 CRS 的更广泛见解。该 SoK 综述系统性地总结了 autonomous CRS(即 agent 用于漏洞发现与修复)在 AIxCC 中的设计、架构与实践经验。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2602.07666v4" target="_blank">https://arxiv.org/abs/2602.07666v4</a></div></div></article></div><div class="conf-group"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">CCS 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">ACM 计算机与通信安全会议</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 11 月 · Cycle 1 通知(7/17),完整列表待公开</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">10</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">PBFuzz: Agentic Directed Fuzzing for PoV Generation</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">PBFuzz:面向 PoV 生成的智能体化导向模糊测试</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">CCS 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Haochen Zeng, Andrew Bao, Jiajun Cheng, Chengyu Song</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国加州大学河滨分校 University of California, Riverside）</span> · <span class="affil" style="color: var(--accent2);">（美国明尼苏达大学双城分校 University of Minnesota, Twin Cities）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>PoV(Proof-of-Vulnerability)输入生成是软件安全中的关键任务,支持路径生成与验证等下游应用。生成 PoV 输入需要求解两类约束:(1)到达漏洞代码位置的可达性约束;(2)激活目标漏洞的触发约束。现有方法(包括导向灰盒模糊测试和 LLM 辅助模糊测试)难以高效满足这些约束。本工作提出一种模仿人类专家的智能体化方法:人类分析师迭代研究代码以提取语义可达性和触发约束,形成关于 PoV 触发策略的假设,将其编码为测试输入,并利用调试反馈精炼其理解。我们用 PBFuzz——一个智能体化导向模糊测试框架——自动化此过程。PBFuzz 解决了 agentic PoV 生成中的四个挑战:自主代码推理提取语义约束、自定义程序分析工具进行目标推理、持久记忆避免假设漂移、以及基于属性的测试以高效求解约束同时保留输入结构。在 Magma 基准上的实验显示,PBFuzz 触发了 57 个漏洞(超过所有基线),并独占触发了 17 个其他方法无法触发的漏洞;PBFuzz 在每个目标 30 分钟预算内实现,而传统方法用 24 小时;中位暴露时间为 339 秒(PBFuzz)对比 8680 秒(AFL++ with CmpLog),效率提升 25.6 倍。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/html/2512.04611v2" target="_blank">https://arxiv.org/html/2512.04611v2</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">代码</span><a href="https://github.com/sgzeng/pbfuzz" target="_blank">https://github.com/sgzeng/pbfuzz</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">11</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">BACAgent: LLM-Powered Detection of Broken-Access-Control Vulnerabilities in Web Applications</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">BACAgent:基于 LLM 的 Web 应用越权访问控制漏洞检测</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">CCS 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Fengyu Liu, Yuan Zhang, Zheng Lou, Tian Chen, Youkun Shi, Jiarun Dai, Enhao Li, Guangyu Zhou, Zhongfu Su, Zequn Fang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（中国复旦大学 Fudan University）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>BACAgent 是一种基于 LLM 的智能体系统,用于自动检测 Web 应用中的 Broken-Access-Control(BAC,越权访问控制)漏洞。BAC 是 Web 应用中最常见且最具危害的漏洞类型之一,但传统静态/动态分析工具难以有效识别此类涉及业务逻辑的漏洞。BACAgent 利用 LLM 智能体对 Web 应用的权限检查逻辑进行语义理解、跨服务跟踪和漏洞推理,实现对 BAC 漏洞的自动化检测。该方法克服了传统工具在理解复杂业务逻辑和跨服务权限传递方面的局限,展示了智能体在逻辑漏洞检测领域的应用潜力。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://lfysec.github.io/" target="_blank">https://lfysec.github.io/</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">12</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">The Illusion of Rust Safety: Detecting Modular Unsafe Functions with LLMs (COIN)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Rust 安全的幻象:使用 LLM 检测模块化不安全函数</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">CCS 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Xiang Cheng, Fan Sang, Yibin Yang, Hang Zhang, Sangdon Park, Xiaokuan Zhang, Taesoo Kim</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国佐治亚理工学院 Georgia Institute of Technology）</span> · <span class="affil" style="color: var(--accent2);">（其他合作机构）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>内存安全漏洞约占已报告安全 bug 的 70%,促使了 Rust 的采用——这是一种通过编译时检查防止此类问题的系统编程语言。虽然 Rust 将代码分为 safe 和 unsafe 区域,但我们识别出一类此前未被充分研究的漏洞:Rust 代码中的模块化不安全函数(MUFs)。这些函数对编译器看似安全,但在整个模块的 API 上下文中实际上是不安全的,可在不使用 unsafe 关键字的情况下导致未定义行为,造成误导性的安全感和 Rust 中的 unsound 问题。本文首次对 MUFs 进行系统研究,从经验分析中定义了五个根因类别。我们开发了 COIN,一种使用微调大语言模型进行漏洞识别和 PoC 生成的检测系统。COIN 在标注评估集上达到 63.7% 精度和 80.4% 召回率,超越现有工具。在对随机抽样的 Rust crate 的评估中,COIN 识别出 26 个未知 MUFs,其中 14 个获开发者确认,10 个被分配 CVE 编号。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://sangdon.github.io/publication/2026-05-16-the-illusion-of-rust-safety-detecting-modular-unsafe-functions-wi/" target="_blank">https://sangdon.github.io/publication/2026-05-16-the-illusion-of-rust-safety-detecting-modular-unsafe-functions-wi/</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">13</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Kumushi: Root-Cause-Driven Automated Vulnerability Repair</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Kumushi:根因驱动的自动化漏洞修复</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">CCS 2026(待确认)</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Hulin Wang, Zion Leonahenahe Basque, Jie Hu, Ati Priya Bajaj, Yibo Liu, Samuel Zhu, Giorgi Kobakhia, Nikhil Chapre, Will Rosenberg, Siddharth Mishra, Aditya Maheshbhai Gabani, Moritz Schloegel, Adam Doupé, Yan Shoshitaishvili, Ruoyu Wang, Tiffany Bao</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国亚利桑那州立大学 Arizona State University）</span> · <span class="affil" style="color: var(--accent2);">（德国 CISPA 亥姆霍兹信息安全中心）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>近期的 LLM 系统使自动化漏洞修复变得日益实用,但仍有两大挑战:第一,没有关于 bug 起源的强信号,修复智能体(repair agent)会倾向于浅层编辑——虽然消除了观察到的失败但未解决底层缺陷。第二,找到 bug 的根因很困难:即使是熟悉代码库的开发者也经常修复症状而非根因,LLM 智能体在更嘈杂的上下文和更弱的程序理解下也不例外。本文提出 Kumushi,一种根因驱动的修复智能体(patching agent),通过将多样化的动态故障定位与证据加权排序相结合,将 LLM 的注意力集中在与缺陷最相关的代码上。为严格衡量 Kumushi 是否产生真正更好的补丁,论文还引入了一种两级补丁质量指标,将自动化 oracle 验证与结构化专家评估相结合。在 178 个 C/C++ 漏洞上的评估中,Kumushi 在自动化评估下显著优于先前的专用修复智能体,并匹敌前沿的商业编码智能体(OpenAI Codex)。专家评估显示:在两者都产生可信补丁的 144 个 bug 中,人类专家盲评在 68 个(47%)上更偏好 Kumushi,39 个(27%)偏好 Codex,37 个(26%)认为等同。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/html/2605.04251v1" target="_blank">https://arxiv.org/html/2605.04251v1</a></div></div><div class="notice notice-warn" style="margin-top: 0.8rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: rgba(161, 98, 7, 0.08);border-left-color: rgb(161, 98, 7);border-radius: 0px 6px 6px 0px;font-size: 0.82rem;color: var(--ink);"><strong style="color: rgb(161, 98, 7);">归属说明:</strong> arXiv 论文脚注写&#34;ACM Conference on Computer and Communications Security&#34;,强烈暗示为 CCS 2026 论文,但未在作者公开页面中明确列出年份,请以最终 CCS 2026 官方论文列表为准。</div></article></div></div></p><p style="padding-top: 2.5rem;padding-bottom: 2.5rem;border-top: 1px solid var(--rule);color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="container" style="margin-right: auto;margin-left: auto;padding-right: 1.5rem;padding-left: 1.5rem;max-width: 980px;"><div class="section-head" style="margin-bottom: 2.5rem;"><div class="kicker" style="margin-bottom: 0.6rem;font-family: JetBrainsMono, monospace;font-size: 0.78rem;letter-spacing: 0.15em;text-transform: uppercase;color: var(--accent);">PART 2 · 软工与安全测试顶会</div><h2 style="margin-bottom: 0.6rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.5rem;font-weight: 700;color: var(--ink);line-height: 1.3;">软件工程顶会与安全测试顶会(ICSE / FSE / ASE / ISSTA)</h2><p class="lead" style="color: var(--muted);font-size: 1rem;max-width: 720px;">软件工程与安全测试顶会是自动化程序修复与漏洞检测的重要发表渠道。2026 年 ICSE 与 FSE 共检索到 4 篇相关论文,ASE 与 ISSTA 因 proceedings 尚未公开暂未找到明确归属论文。</p></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">ICSE 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">国际软件工程会议</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 4 月 · 已完整公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">14</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Fixing Security Vulnerabilities with Agentic AI in OSS-Fuzz (CodeRover-S)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">使用智能体 AI 修复 OSS-Fuzz 中的安全漏洞</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICSE 2026</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Yuntong Zhang, Jiawei Wang, Dominic Berzin, Martin Mirchev, Abhik Roychoudhury</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（新加坡国立大学）</span> · <span class="affil" style="color: var(--accent2);">（新加坡国立大学）</span> · <span class="affil" style="color: var(--accent2);">（新加坡国立大学）</span> · <span class="affil" style="color: var(--accent2);">（SonarSource）</span> · <span class="affil" style="color: var(--accent2);">（新加坡国立大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>关键开源软件系统通过长期的模糊测试活动进行大量验证。模糊测试对程序输入域进行有偏随机搜索以发现使软件崩溃的输入。Google 的 OSS-Fuzz 是持续验证开源系统的最重要基础设施,已识别超过 13,000 个漏洞,但这些漏洞往往仍未被修补,因为修补在实践中通常是手动的。本研究探索使用大语言模型(LLM)智能体进行自动化漏洞修复,据我们所知这是首次对 OSS-Fuzz 上 LLM 辅助安全补丁进行系统研究。我们将通常从 issue 描述修复 bug 的 AutoCodeRover 智能体适配到安全领域,提出 CodeRover-S。该智能体不使用 issue 文本,而是通过执行漏洞利用输入提取与漏洞相关的代码元素,并用静态类型信息增强补丁生成。在历史漏洞基准上,智能体在 61%–72% 的案例中生成了合理补丁;在 OSS-Fuzz 报告的真实未修补漏洞上,智能体生成了 73.3% 的合理补丁,多个补丁已被合并到广泛使用的开源项目中。这证明了 LLM 智能体自动化漏洞修复的实用性,以及从检测到修复的端到端软件保护周期的可行性。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://dl.acm.org/doi/10.1145/3786583.3786880" target="_blank">https://dl.acm.org/doi/10.1145/3786583.3786880</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">15</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents (VulTrial)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">审判开始:基于 LLM 智能体的法庭模拟漏洞检测方法</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICSE 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Ratnadira Widyasari, Martin Weyssow, Ivana Clairine Irsan, Han Wei Ang, Frank Liauw, Eng Lieh Ouh, Lwin Khin Shar, Hong Jin Kang, David Lo</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（新加坡管理大学）</span> · <span class="affil" style="color: var(--accent2);">（新加坡政府科技局 GovTech）</span> · <span class="affil" style="color: var(--accent2);">（悉尼大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>检测源代码中的漏洞仍是一项关键且具有挑战性的任务,尤其是当良性函数与漏洞函数高度相似时。本文提出 VulTrial,一个受法庭审判启发的多智能体框架,用于增强自动化漏洞检测。它采用四个角色特定的智能体:安全研究员、代码作者、主持人(moderator)和评审委员会(review board)。通过使用 GPT-3.5 和 GPT-4o 的广泛实验,证明 VulTrial 优于单智能体和多智能体基线。使用 GPT-4o 时,VulTrial 将正确标注的样本对数比各自基线提高了 41 和 37(约翻倍)。此外,使用少量数据(50 对样本)对 VulTrial 进行角色特定指令微调,可将正确标注对数提高 56 和 52 个点。研究还分析了增加智能体交互次数对 VulTrial 整体性能的影响,发现将 VulTrial 应用于像 GPT-3.5 这样低成本的模型,可使其性能超过单智能体设置下的 GPT-4o,且总成本更低。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://conf.researchr.org/details/icse-2026/icse-2026-research-track/214/Let-the-Trial-Begin-A-Mock-Court-Approach-to-Vulnerability-Detection-using-LLM-Based" target="_blank">https://conf.researchr.org/details/icse-2026/icse-2026-research-track/214/Let-the-Trial-Begin-A-Mock-Court-Approach-to-Vulnerability-Detection-using-LLM-Based</a></div></div></article></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">FSE 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">欧洲软件工程会议/软件工程基础研讨会</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 7 月 · 已公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">16</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">ChainDelta: Automatic Patch-based Exploit Generation for Ethereum with Fuzzing Agents</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">ChainDelta:基于模糊测试智能体的以太坊补丁驱动漏洞利用自动生成</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">FSE 2026</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Mingxi Ye, Yuhong Nan, Zhijie Zhong, Jianzhong Su, Xingwei Lin, Peilin Zheng, Zibin Zheng</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（中山大学）</span> · <span class="affil" style="color: var(--accent2);">（中山大学软件工程学院）</span> · <span class="affil" style="color: var(--accent2);">（浙江大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>鉴于以太坊的关键性,利用已修补但尚未广泛部署的 1-day 漏洞至关重要。自动补丁驱动漏洞利用生成(APEG)是一项有前景的技术,可帮助开发者理解根因、验证下游 fork 的修复并检测不完整补丁。然而现有漏洞利用生成工具在以太坊上效果不佳,存在三大挑战:(1)导航补丁中隐藏的复杂跨语言利用路径;(2)合成复杂的有状态环境配置;(3)处理区块链节点间导致误报的非确定性不一致。本文提出 ChainDelta,一个由大语言模型驱动的模糊测试智能体框架,基于以太坊安全补丁自动生成漏洞利用。ChainDelta 包含三个核心模块:定向模糊器利用调用图分析引导测试走向补丁信息指向的漏洞代码;基于智能体的环境模糊器作为专家自动设置触发漏洞所需的区块链状态;状态感知消毒器在监控区块链瞬态状态的同时执行差异分析,以区分真实不一致与良性非确定性。在包含真实漏洞补丁的多样基准上评估,ChainDelta 以 70% 成功率生成漏洞利用,误报率仅 12.5%。真实审计活动还发现了 4 个此前未披露的漏洞并获得赏金。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://conf.researchr.org/details/fse-2026/fse-2026-research-papers/98/ChainDelta-Automatic-Patch-based-Exploit-Generation-for-Ethereum-with-Fuzzing-Agents" target="_blank">https://conf.researchr.org/details/fse-2026/fse-2026-research-papers/98/ChainDelta-Automatic-Patch-based-Exploit-Generation-for-Ethereum-with-Fuzzing-Agents</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">17</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">CodeCureAgent:静态分析告警的自动分类与修复</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">FSE 2026</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">告警验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Pascal Joos, Islem Bouzenia, Michael Pradel</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（CISPA 亥姆霍兹信息安全中心,德国）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>静态分析工具广泛用于检测 bug、漏洞和代码异味。开发者通常必须手动解决这些告警。本文提出 CodeCureAgent,利用基于 LLM 的智能体自动分析、分类和修复静态分析告警。与以往遵循预定算法的方法不同,CodeCureAgent 采用智能体框架,迭代调用工具从代码库中收集额外信息(如代码搜索)并编辑代码库以解决告警;它能检测并抑制误报,同时修复识别出的真正告警。采用三步启发式审批补丁:(1)构建项目;(2)验证告警消失且不引入新告警;(3)运行测试套件。在 106 个 Java 项目的 1,000 条 SonarQube 告警(覆盖 291 条规则)上评估,方法为 96.8% 的告警生成合理修复,分别比基线高 30.7% 和 29.2%;手动检查 291 例显示正确修复率 86.3%。注:本文目标为静态分析告警(含安全漏洞,但不限于漏洞),与&#34;漏洞修复&#34;部分相关。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://doi.org/10.1145/3808140" target="_blank">https://doi.org/10.1145/3808140</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">预印本</span><a href="https://arxiv.org/abs/2509.11787" target="_blank">https://arxiv.org/abs/2509.11787</a></div></div></article></div><div class="conf-group"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">ASE / ISSTA</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">自动化软件工程 / 软件测试与分析研讨会</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 10 月 · proceedings 未公开</span></div><div class="notice" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: var(--accent2-soft);border-left: 4px solid var(--accent2);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: var(--accent2);">状态说明:</strong> ASE 2026(10 月 12-16 日,德国慕尼黑)录用通知于 2026 年 7 月 10 日刚发出,ISSTA 2026(10 月 3-9 日,美国奥克兰)录用通知于 6 月 25 日发出,截至检索日(7 月 19 日)两会议的官方完整 proceedings 尚未公开。已公开的少量论文中未找到明确以&#34;agent 用于漏洞挖掘/验证/利用/修复&#34;为主题的论文。<strong style="color: var(--accent2);">建议在 proceedings 正式公开后复查</strong>,本报告第七节列出了多篇时间点与 ASE/ISSTA 通知期吻合的相关 arXiv 预印本供参考。</div></div></div></p><p style="padding-top: 2.5rem;padding-bottom: 2.5rem;border-top: 1px solid var(--rule);color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="container" style="margin-right: auto;margin-left: auto;padding-right: 1.5rem;padding-left: 1.5rem;max-width: 980px;"><div class="section-head" style="margin-bottom: 2.5rem;"><div class="kicker" style="margin-bottom: 0.6rem;font-family: JetBrainsMono, monospace;font-size: 0.78rem;letter-spacing: 0.15em;text-transform: uppercase;color: var(--accent);">PART 3 · AI 顶会</div><h2 style="margin-bottom: 0.6rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.5rem;font-weight: 700;color: var(--ink);line-height: 1.3;">人工智能顶会(ICLR / ICML / AAAI / 其他)</h2><p class="lead" style="color: var(--muted);font-size: 1rem;max-width: 720px;">AI 顶会中相关论文主要集中在 ICLR 2026(3 篇主会)与 ICML 2026,以&#34;评估 AI 智能体网络安全能力&#34;与&#34;红队测试&#34;为主线。AAAI 与 CVPR 主会未找到严格符合主题的论文,NeurIPS 2026 与 IJCAI 2026 因时间原因无法完整检索。</p></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">ICLR 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">国际学习表征会议</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 5 月 · 已完整公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">18</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">CyberGym: Evaluating AI Agents&#39; Real-World Cybersecurity Capabilities at Scale</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">CyberGym:大规模评估 AI 智能体真实世界网络安全能力</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICLR 2026(Oral)</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（加州大学伯克利分校 UC Berkeley,Dawn Song 团队）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>AI 智能体具有重塑网络安全的巨大潜力,因此对其能力进行充分评估至关重要。然而现有评估多基于小规模基准并仅测量静态结果,无法涵盖真实世界安全挑战的动态全貌。为此本文提出 CyberGym,一个大规模基准,涵盖 188 个开源软件项目中的 1,507 个真实世界漏洞。该基准可适配不同的漏洞分析设置,主要任务是在仅给出漏洞文本描述和对应代码库的情况下,让智能体生成可复现该漏洞的概念验证(PoC)测试。大规模评估表明 CyberGym 能有效区分不同智能体和模型的网络安全能力,即使表现最好的组合也仅达到约 20% 的成功率,说明 CyberGym 整体难度很高。除静态基准测试外,CyberGym 还促成了 34 个零日漏洞和 18 个历史不完整补丁的发现。这些结果表明 CyberGym 不仅是衡量 AI 在网络安全领域进展的稳健基准,也是产生直接真实世界安全影响的平台。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://openreview.net/forum?id=2YvbLQEdYt" target="_blank">https://openreview.net/forum?id=2YvbLQEdYt</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">19</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">HackWorld:评估计算机操作型智能体在利用 Web 应用漏洞方面的能力</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICLR 2026</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Xiaoxue Ren, Penghao Jiang, Jiaojiao Jiang, Kaixin Li, Zhiyong Huang, Xiaoning Du, Terry Yue Zhuo, Zhenchang Xing, Jiamou Sun</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（浙江大学软件学院）</span> · <span class="affil" style="color: var(--accent2);">（澳大利亚新南威尔士大学 UNSW）</span> · <span class="affil" style="color: var(--accent2);">（新加坡国立大学 NUS）</span> · <span class="affil" style="color: var(--accent2);">（澳大利亚莫纳什大学 Monash）</span> · <span class="affil" style="color: var(--accent2);">（澳大利亚联邦科工组织 CSIRO&#39;s Data61）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>Web 应用因作为关键服务和敏感数据的入口而成为网络攻击的主要目标。传统渗透测试昂贵且需要专门知识。虽然语言模型智能体在某些网络安全任务上已展现潜力,但现代 Web 应用需要通过复杂用户界面、动态内容渲染和多步交互工作流的视觉理解,这只有计算机操作型智能体(CUA)能处理。本文提出 HackWorld,首个系统性评估 CUA 通过可视化交互发现和利用 Web 应用漏洞能力的框架。与使用净化环境的现有基准不同,HackWorld 让 CUA 面对 36 个精心策划的应用(覆盖 11 个框架和 7 种语言),包含注入缺陷、身份验证绕过、不安全输入处理等真实漏洞。该框架采用夺旗赛(CTF)方法直接评估 CUA 在导航复杂 Web 界面时发现和利用这些漏洞的能力。对最先进 CUA 的评估显示,其漏洞利用率低于 12%,在多步攻击规划和安全工具使用方面存在困难。结果揭示了 CUA 在操作含漏洞 Web 应用时网络安全技能的局限性。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://openreview.net/pdf?id=nLfZPoJbO7" target="_blank">https://openreview.net/pdf?id=nLfZPoJbO7</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">20</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">RedCodeAgent:针对多样化代码智能体的自动化红队智能体</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICLR 2026(Poster)</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Chengquan Guo, Chulin Xie, Yu Yang, Zhaorun Chen, Zinan Lin, Xander Davies, Yarin Gal, Dawn Song, Bo Li</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（加州大学伯克利分校 UC Berkeley）</span> · <span class="affil" style="color: var(--accent2);">（微软研究院）</span> · <span class="affil" style="color: var(--accent2);">（芝加哥大学）</span> · <span class="affil" style="color: var(--accent2);">（牛津大学）</span> · <span class="affil" style="color: var(--accent2);">（伊利诺伊大学香槟分校 UIUC）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>代码智能体因强大的代码生成能力和与代码解释器的集成而被广泛采用。然而这些进展也带来了关键的安全风险。现有静态安全基准和红队工具不足以识别新兴的真实风险场景,无法覆盖诸如不同越狱工具组合效应等边界条件。本文提出 RedCodeAgent,首个专门设计用于系统性发现多样化代码智能体中漏洞的自动化红队智能体。借助自适应记忆模块,RedCodeAgent 能利用已有的越狱知识,针对给定输入查询动态选择最有效的红队工具及其组合,从而识别可能被忽视的漏洞。为可靠评估,作者开发了模拟沙盒环境来评估代码智能体的执行结果,缓解仅依赖静态代码的 LLM 评判偏差。在多个最先进代码智能体、多样化风险场景和多种编程语言上的广泛评估表明,RedCodeAgent 持续优于现有红队方法,攻击成功率更高、拒绝率更低、效率更高。作者还在真实世界代码助手(如 Cursor 和 Codeium)上验证了 RedCodeAgent,暴露了此前未识别的安全风险。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://openreview.net/forum?id=IyIaAOihmZ" target="_blank">https://openreview.net/forum?id=IyIaAOihmZ</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">代码</span><a href="https://github.com/1mocat/RedCodeAgent" target="_blank">https://github.com/1mocat/RedCodeAgent</a></div></div></article><div class="notice" style="margin-top: 1rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: var(--accent2-soft);border-left: 4px solid var(--accent2);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: var(--accent2);">ICLR 2026 研讨会论文(AIWILD Workshop,非主会):</strong></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">21</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">ZeroDayBench:评估 LLM 智能体在未见零日漏洞上的网络防御能力</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-workshop" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent2);color: rgb(0, 0, 0);">ICLR 2026 研讨会</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Nancy Lau, Louis Sloot, Jyoutir Raj, Evan Harris, Giuseppe Marco Boscardin, Dan Zhao, Dylan Bowman, Mario Brajkovski, Jaideep Singh Chawla</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（作者具体机构未在公开页面明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>大语言模型(LLM)越来越多地被部署为软件工程智能体,自主参与代码仓库贡献。这些智能体的一大优势在于能在其监管的代码库中发现并修复安全漏洞。为评估智能体在此领域的能力,本文提出 ZeroDayBench,一个让 LLM 智能体在开源代码库中发现并修复 22 个新型严重漏洞的基准。研究聚焦于三种流行的前沿智能体 LLM:GPT-5.2、Claude Sonnet 4.5 和 Grok 4.1。结果显示前沿 LLM 尚不能自主完成这些任务,并观察到一些行为模式,提示了这些模型在主动网络防御领域的改进方向。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://openreview.net/forum?id=ppQy3tSEmH" target="_blank">https://openreview.net/forum?id=ppQy3tSEmH</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">22</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">When Fuzzing Becomes Agentic: Semantic State Exploration in the Wild (ASA-Fuzz)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">当模糊测试变得智能体化:真实世界中的语义状态探索</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-workshop" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent2);color: rgb(0, 0, 0);">ICLR 2026 研讨会</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Andrew Yin, Zhaoling Chen, Qian Zhang, Heng Yin</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国加州大学河滨分校 UC Riverside,Heng Yin 团队）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>智能体 AI 系统越来越多地穿越有状态、受规则约束的软件程序,其中有意义的行为只有在经过一系列有效且上下文感知的交互后才显现。现有评估方法多关注结果,对智能体是否到达了新颖的内部状态缺乏可见性。传统自动化测试往往假设无状态环境并生成违反一致性约束的无效交互。本文提出一种面向有状态程序的闭环智能体测试框架 ASA-Fuzz,将软件模糊测试视为一个智能体决策问题。该框架(1)推断并插桩与状态相关的运行时信号,以提供语义上的进展概念;(2)合成上下文感知的算子,根据观察到的状态生成有效的多步交互;(3)在探索进入平台期时进行自适应调整。在基于规则的有状态国际象棋环境案例研究中,ASA-Fuzz 发现的独特棋盘状态比传统方法多两个数量级,且平均有效移动深度比工业标准(AFL++)、状态感知(SGFuzz)和 LLM 驱动(Fuzz4All)基线高出 3.5 至 12 倍。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://openreview.net/forum?id=ztjVY9GB1t" target="_blank">https://openreview.net/forum?id=ztjVY9GB1t</a></div></div></article></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">ICML 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">国际机器学习会议</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 7 月 · 已公布</span></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">23</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks (SusVibes)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">在真实世界任务中基准测试智能体生成代码的漏洞性</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-conf" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent);color: rgb(0, 0, 0);">ICML 2026</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘(评估)</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, Lei Li</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（卡内基梅隆大学 CMU）</span> · <span class="affil" style="color: var(--accent2);">（哥伦比亚大学 Columbia University）</span> · <span class="affil" style="color: var(--accent2);">（约翰斯·霍普金斯大学 Johns Hopkins）</span> · <span class="affil" style="color: var(--accent2);">（HydroX AI）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>该工作研究&#34;氛围编程(vibe coding)&#34;是否安全。作者构建了一个基准,用于评估在真实世界任务中由 LLM 智能体生成的代码是否存在安全漏洞。任务描述只包含期望功能,不含实现细节或安全指引。研究发现:在通过功能测试的智能体解决方案中,79.3% 是不安全的。前沿 LLM 和主流智能体在安全性方面表现极差,初步的安全尝试也不够,编码智能体需要更好的安全策略。注:该论文是评估&#34;智能体生成代码&#34;的漏洞性,智能体本身是被评估对象,并非直接用 agent 去做漏洞挖掘/修复,属边界相关论文。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://icml.cc/media/icml-2026/Slides/61427.pdf" target="_blank">https://icml.cc/media/icml-2026/Slides/61427.pdf</a></div></div></article><div class="notice" style="margin-top: 1rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: var(--accent2-soft);border-left: 4px solid var(--accent2);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: var(--accent2);">ICML 2026 研讨会论文(AIWILD Workshop,非主会):</strong></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">24</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">红队测试智能体执行上下文:在 OpenClaw 上的开放世界安全评估</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-workshop" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent2);color: rgb(0, 0, 0);">ICML 2026 研讨会</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">(作者列表未在检索片段中完整列出)</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未明确）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>本工作研究 OpenClaw 这类结合 LLM 推理与系统级行为的执行中心型智能体系统的安全边界。OpenClaw 将大语言模型推理与系统级动作和面向用户的工作流结合,实现异构数字环境中的端到端任务完成,这扩展了安全边界。作者在开放世界设置下对 OpenClaw 的执行上下文进行红队安全评估,旨在发现此类智能体执行环境中的安全漏洞。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/pdf/2605.11047.pdf" target="_blank">https://arxiv.org/pdf/2605.11047.pdf</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">25</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Smarter Saboteurs, Better Fixers: Scaling &amp; Security in Linear Multi-Agent Workflows</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">更聪明的破坏者,更好的修复者:线性多智能体工作流中的扩展性与安全性</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-workshop" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent2);color: rgb(0, 0, 0);">ICML 2026 研讨会</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">(作者列表未在检索片段中完整列出)</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未明确）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>该工作研究多智能体系统(MAS)在内部威胁下的安全性和弹性。作者设计了 QA+Fixer 拓扑,其中 QA 工程师审查代码并发出&#34;无问题&#34;或&#34;发现问题&#34;状态,若发现问题则交由 Fixer 工程师生成最终修复后的代码。研究比较了控制场景和恶意工程师场景,探讨线性多智能体工作流中的破坏与修复,展示多智能体工作流中修复智能体对抗恶意注入的能力。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/pdf/2606.12709" target="_blank">https://arxiv.org/pdf/2606.12709</a></div></div></article></div><div class="conf-group" style="margin-bottom: 3rem;"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">AAAI 2026</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">AAAI 人工智能会议</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">2026 年 1 月 · 主会无相关论文</span></div><div class="notice notice-warn" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: rgba(161, 98, 7, 0.08);border-left-color: rgb(161, 98, 7);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: rgb(161, 98, 7);">主会情况:</strong> 经搜索 AAAI 2026 主会(2026 年 1 月 20-27 日,新加坡)论文,未找到以&#34;agent 用于漏洞挖掘/验证/利用/修复&#34;为主题的主会论文。检索到的 AAAI 2026 相关论文多为&#34;攻击智能体&#34;或&#34;智能体安全&#34;主题(如越狱/对抗攻击论文),不符合本任务&#34;agent 作为漏洞工作的执行者&#34;的界定。</div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">26</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Reflection-Driven Control for Trustworthy Code Agents</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">面向可信代码智能体的反思驱动控制</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-workshop" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: var(--accent2);color: rgb(0, 0, 0);">AAAI 2026 研讨会</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复(预防性)</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">(arXiv 编号 2512.21354,作者未在检索片段完整列出)</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未在公开页面明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>本文提出反思驱动控制(Reflection-Driven Control)模块,将&#34;自我反思&#34;从一种事后补丁提升为智能体推理过程中的一等控制回路。该模块包含三个组件:轻量级自检查器、证据驱动修复和反思性记忆库。在安全代码生成任务上,该模块显著提升了代码安全率。整体方法采用 Plan-Reflect-Verify 三阶段框架,通过自我检查的轻量级预过滤、动态记忆 RAG 驱动的修复以及编译验证反馈回路实现持续安全控制。实验显示,在 8 个 CWE 漏洞类别上,安全率平均提升约 8.1 个百分点,而功能通过率基本保持稳定。该论文侧重&#34;让智能体生成更安全的代码&#34;,属预防性安全,与&#34;漏洞修复&#34;主题部分相关。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://trustagenticai.github.io/AAAI2026/AAAI-Workshop/43.pdf" target="_blank">https://trustagenticai.github.io/AAAI2026/AAAI-Workshop/43.pdf</a></div><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">预印本</span><a href="https://arxiv.org/abs/2512.21354" target="_blank">https://arxiv.org/abs/2512.21354</a></div></div></article></div><div class="conf-group"><div class="conf-header" style="margin-bottom: 1.5rem;padding-bottom: 0.8rem;display: flex;align-items: center;gap: 0.8rem;border-bottom: 2px solid var(--accent);flex-wrap: wrap;"><span class="conf-badge" style="padding: 0.3rem 0.7rem;display: inline-block;background: var(--accent);color: rgb(0, 0, 0);font-family: JetBrainsMono, monospace;font-size: 0.75rem;font-weight: 700;border-radius: 3px;letter-spacing: 0.04em;">CVPR / IJCAI / NeurIPS</span><span class="conf-title" style="font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.3rem;font-weight: 700;color: var(--ink);">其他 AI 顶会状态</span><span class="conf-note" style="font-size: 0.85rem;color: var(--muted);width: 412px;">未找到 / 未召开</span></div><div class="notice notice-warn" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: rgba(161, 98, 7, 0.08);border-left-color: rgb(161, 98, 7);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: rgb(161, 98, 7);">CVPR 2026:</strong> 经搜索未找到以&#34;agent 用于漏洞挖掘/验证/利用/修复&#34;为主题的论文。CVPR 2026 中相关论文均为&#34;智能体作为被攻击目标&#34;的主题(如 AGENTSAFE、AdapAction 后门攻击),不符合本任务要求。</div><div class="notice" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: var(--accent2-soft);border-left: 4px solid var(--accent2);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: var(--accent2);">IJCAI-ECAI 2026:</strong> 将于 2026 年 8 月 15-21 日在德国不来梅召开,共接收 990 篇论文。由于会议尚未召开,完整论文列表尚未完全公开,未找到明确相关论文。</div><div class="notice" style="margin-top: 1.5rem;margin-bottom: 1.5rem;padding: 1rem 1.2rem;background: var(--accent2-soft);border-left: 4px solid var(--accent2);border-radius: 0px 6px 6px 0px;font-size: 0.9rem;color: var(--ink);"><strong style="color: var(--accent2);">NeurIPS 2026:</strong> 会议将于 2026 年 12 月 6-14 日召开,作者通知为 2026 年 9 月 24 日。截至检索日论文尚未公布,无法检索。建议在 9 月 24 日通知后再次查询。</div></div></div></p><p style="padding-top: 2.5rem;padding-bottom: 2.5rem;border-top: 1px solid var(--rule);color: rgb(26, 35, 50);font-family: WorkSans, -apple-system, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, &#34;Hiragino Sans GB&#34;, &#34;Noto Sans CJK SC&#34;, sans-serif;font-size: 16px;letter-spacing: normal;text-align: start;"><div class="container" style="margin-right: auto;margin-left: auto;padding-right: 1.5rem;padding-left: 1.5rem;max-width: 980px;"><div class="section-head" style="margin-bottom: 2.5rem;"><div class="kicker" style="margin-bottom: 0.6rem;font-family: JetBrainsMono, monospace;font-size: 0.78rem;letter-spacing: 0.15em;text-transform: uppercase;color: var(--accent);">PART 4 · 预印本参考</div><h2 style="margin-bottom: 0.6rem;font-family: InstrumentSans, &#34;PingFang SC&#34;, &#34;Microsoft YaHei&#34;, sans-serif;font-size: 1.5rem;font-weight: 700;color: var(--ink);line-height: 1.3;">高度相关的 arXiv 预印本(会议归属未确认)</h2><p class="lead" style="color: var(--muted);font-size: 1rem;max-width: 720px;">以下论文与主题高度相关,但其 arXiv 预印本未标注目标会议归属,不能确认已被 2026 年目标会议接收。部分论文时间点(6 月底-7 月初)与 ASE 2026(7 月 10 日通知)/ISSTA 2026(6 月 25 日通知)吻合,可能已录用但未公开。列出供参考。</p></div><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P1</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Synthesizing Multi-Agent Harnesses for Vulnerability Discovery (AgentFlow)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">面向漏洞发现的多智能体运行时框架合成</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Hanzhi Liu, Chaofan Shou, Xiaonan Liu, Hongbo Wen, Yanju Chen, Ryan Jingyang Fang, Yu Feng</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（美国加州大学圣巴巴拉分校 UC Santa Barbara,SECAI 实验室,Yu Feng 团队）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>LLM 智能体已开始发现人类审计员和自动化模糊测试工具数十年未察觉的真实安全漏洞,前提是分析者能够构建并插桩目标代码。实际中此类工作通常由多个智能体通过&#34;运行时框架(harness)&#34;协同完成。当语言模型固定、仅调整 harness 时,在公开智能体基准上任务成功率仍可相差数倍;然而目前大多数 harness 仍需手工编写。本文提出 AgentFlow,通过一个具有类型约束的图结构 DSL 以及一个反馈驱动的外层优化循环,联合覆盖智能体角色、提示词、工具、通信拓扑和协同协议。在 TerminalBench-2 上达到 84.3%,并在 Google Chrome 中发现了 10 个此前未知的零日漏洞,包括 2 个严重级别的沙箱逃逸漏洞(CVE-2026-5280 和 CVE-2026-6297)。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2604.20801" target="_blank">https://arxiv.org/abs/2604.20801</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P2</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Code-Augur: Agentic Vulnerability Detection via Specification Inference</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Code-Augur:通过规范推断的智能体化漏洞检测</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Zhengxiong Luo, Mehtab Zafar, Dylan Wolff, Abhik Roychoudhury</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（新加坡国立大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>提出安全规范优先的范式,智能体将安全假设显式化为源码内断言,并用引导式模糊器持续证伪,形成&#34;推理-证伪-精化&#34;闭环。在 AIxCC 和 OSV 基准上比 Claude Code、Atlantis 多发现 34%–370% 的 bug,并发现 22 个新漏洞(16 个已修复/确认)。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2606.18619" target="_blank">https://arxiv.org/abs/2606.18619</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P3</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">KeaRepair: Knowledge-Enhanced Agentic Vulnerability Repair</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">KeaRepair:知识增强的智能体漏洞修复</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Sicong Cao, Hao Ma, Le Yu, Kangyi Ding, Xiaolei Liu, Terry Yue Zhuo, Bo Wang, Xingwei Lin, Xiaobing Sun, Linzhang Wang, David Lo</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未在 arXiv 明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>提出工具增强的 ReAct 风格智能体,从历史漏洞-补丁对提取多维漏洞知识构建检索库,进行漏洞诊断与知识级检索增强补丁生成,闭环验证。在 55 个可复现 C/C++ 漏洞上修复率达 83.64%。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2607.00820" target="_blank">https://arxiv.org/abs/2607.00820</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P4</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Refploit: Facilitating Exploit Construction via Code-Agent Trajectory Repair</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Refploit:通过代码智能体轨迹修复辅助漏洞利用构造</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-exploit" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(148, 163, 184, 0.18);color: rgb(71, 85, 105);">漏洞利用</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Zirui Chen, Zhipeng Xue, Jiayuan Zhou, Xing Hu, Xin Xia, Xiaohu Yang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（浙江大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>提出基于 LLM 的轨迹恢复框架,从公开漏洞利用参考出发,通过差分执行验证、定位轨迹片段、推导修复约束,辅助 Java 漏洞利用复现。在 172 个利用参考上复现率 80.2%。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2607.01760" target="_blank">https://arxiv.org/abs/2607.01760</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P5</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">FuzzingBrain V2: A Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">FuzzingBrain V2:用于自动漏洞发现与复现的多智能体 LLM 系统</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Ze Sheng, Zhicheng Chen, Qingxiao Xu, Kewen Zhu, Jeff Huang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未在 arXiv 明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>多智能体系统,基于 OSS-Fuzz 实现全自动漏洞分析;提出 Suspicious Point 抽象进行精确漏洞定位;双层模糊增强函数覆盖。AIxCC 2025 数据集检测率 90%,真实部署发现 29 个 0-day(均被确认修复,2 个获 CVE)。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2605.21779" target="_blank">https://arxiv.org/abs/2605.21779</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P6</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Phoenix: Security Is Relative — Training-Free Vulnerability Detection via Multi-Agent Behavioral Contract Synthesis</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">Phoenix:安全是相对的——通过多智能体行为契约合成的免训练漏洞检测</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Yongchao Wang, Zhiqiu Huang</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（机构未在 arXiv 明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>提出训练无关的多智能体框架,通过行为契约合成解决语义歧义,三阶段:语义切片、需求逆向工程(合成 Gherkin 行为规范)、契约判定。PrimeVul Paired 上 F1=0.825,超过 VulTrial(F1=0.563)。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2604.19012" target="_blank">https://arxiv.org/abs/2604.19012</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P7</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">MulVul:基于跨模型提示演进的检索增强多智能体代码漏洞检测</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-mine" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(185, 28, 28, 0.1);color: var(--accent);">漏洞挖掘</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Zihan Wu, Jie Xu, Yun Peng, Chun Yong Chong, Xiaohua Jia</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（作者所属机构未在 arXiv 明确列出）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>提出检索增强多智能体框架 MulVul,采用&#34;粗到细&#34;策略:Router 智能体先预测 top-k 粗粒度类别,再由专门的 Detector 智能体识别具体漏洞类型。关键创新是&#34;跨模型提示演进&#34;机制——生成器 LLM(如 Claude)迭代提出候选提示,由不同的执行器 LLM(如 GPT-4o)验证有效性。在 130 个 CWE 类型上,MulVul 达到 34.79% Macro-F1,比最优基线提升 41.5%。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/pdf/2601.18847" target="_blank">https://arxiv.org/pdf/2601.18847</a></div></div></article><article class="paper" style="margin-bottom: 1.5rem;padding: 1.3rem;background: var(--bg2);border: 1px solid var(--rule);border-radius: 8px;transition: box-shadow 0.2s, border-color 0.2s;"><div class="paper-head" style="margin-bottom: 0.8rem;display: flex;align-items: flex-start;gap: 1rem;"><span class="paper-num" style="margin-top: 0.15rem;padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.85rem;font-weight: 700;color: var(--accent);background: var(--accent-soft);border-radius: 4px;flex-shrink: 0;">P8</span><div class="paper-titles" style="flex: 1 1 0%;"><div class="paper-title-en" style="margin-bottom: 0.2rem;font-family: InstrumentSans, sans-serif;font-size: 1.15rem;font-weight: 700;color: var(--ink);line-height: 1.35;">Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities (LiLCVE)</div><div class="paper-title-cn" style="font-size: 1rem;font-weight: 600;color: var(--accent2);line-height: 1.4;">揭秘与修复 LLM-in-the-Loop 漏洞</div></div></div><div class="paper-tags" style="margin-top: 0.6rem;margin-bottom: 0.6rem;display: flex;flex-wrap: wrap;gap: 0.4rem;"><span class="tag tag-preprint" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgb(161, 98, 7);color: rgb(0, 0, 0);">arXiv 预印本</span><span class="tag tag-repair" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(22, 163, 74, 0.12);color: rgb(21, 128, 61);">漏洞修复</span><span class="tag tag-verify" style="padding: 0.15rem 0.5rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 600;border-radius: 3px;letter-spacing: 0.02em;background: rgba(30, 64, 175, 0.1);color: var(--accent2);">漏洞验证</span></div><div class="paper-authors" style="margin-bottom: 0.4rem;font-size: 0.9rem;color: var(--ink);line-height: 1.6;"><span class="author">Yujie Ma, Jialin Rong, Chenxi Yang, Lili Quan, Jin Wen, Xiaofei Xie, Yongqiang Lyu, Qiang Hu</span></div><div class="paper-affil" style="margin-bottom: 1rem;font-size: 0.82rem;color: var(--muted);"><span class="affil" style="color: var(--accent2);">（天津大学）</span> · <span class="affil" style="color: var(--accent2);">（新加坡管理大学）</span></div><div class="paper-abstract" style="margin-bottom: 1rem;font-size: 0.92rem;color: var(--ink);line-height: 1.75;text-align: justify;"><span class="label" style="margin-bottom: 0.3rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;font-weight: 700;color: var(--accent);letter-spacing: 0.08em;text-transform: uppercase;display: block;">摘要</span>定义 LLM-in-the-Loop (LiL) 漏洞,构建首个 LiL 漏洞数据集 LiLCVE(41 个 LiL + 75 个 LLM 生态漏洞),并用 20 种 agent-模型配置(含 SWE-Agent)评估修复能力,发现 LiL 漏洞比传统漏洞更难修复。</div><div class="paper-foot" style="padding-top: 0.8rem;display: flex;flex-wrap: wrap;gap: 1rem 2rem;border-top: 1px dashed var(--rule);font-size: 0.85rem;"><div class="field" style="display: flex;align-items: flex-start;gap: 0.4rem;"><span class="field-label" style="margin-top: 0.1rem;font-family: JetBrainsMono, monospace;font-size: 0.72rem;color: var(--muted);text-transform: uppercase;letter-spacing: 0.05em;flex-shrink: 0;">下载</span><a href="https://arxiv.org/abs/2605.28893" target="_blank">https://arxiv.org/abs/2605.28893</a></div></div></article></div></p><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=88266ec6&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486128%26idx%3D1%26sn%3D791ba32939688d0a2e21fdf74b02af4b">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Mon, 20 Jul 2026 11:02:00 +0800</pubDate>
    </item>
    <item>
      <title>USENIX Security 2026 — Cycle 1 论文清单与摘要（上）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486124&amp;idx=1&amp;sn=5feac1df49b63f898b824fc3fd66ccd8</link>
      <description></description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span>2026-07-16 20:22</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=93296ba7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sxAzyxGKMHVU9P5EBW6Kne6s9ssA4o6AHibT0NtA0Pj16ficUDgiaZhJD9xOTIG6jSM6qMvZ7CupcW0CUahENOUNwUqjYoLTpWm2U%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">1. Bond: Constraint-Directed Fuzzing for Automated Validation of Taint Analysis Results in Linux-based IoT Firmware</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jiaqian Peng (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院); Puzhuo Liu (蚂蚁集团; 清华大学); Kai Cheng (中国科学院信息工程研究所); Zhaoteng Yan (中国科学院大学网络空间安全学院); Jie Liu (中国科学院信息工程研究所); Chengnian Sun (滑铁卢大学); Hongsong Zhu (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">IoT 设备的固件漏洞构成严重的安全威胁，然而最先进的污点分析工具往往生成大量报告却缺乏充分验证。我们提出 Bond，一种有向模糊测试框架，连接静态污点分析与动态漏洞验证。Bond 通过集成三大类、六种语义类型的约束，引入约束引导的输入变异，从而高效地探索与污点报告相关联的路径。我们在来自 8 家厂商的 19 款 IoT 设备上评估 Bond，覆盖四种最先进污点分析器所产生的 2,776 份污点报告。Bond 成功验证了 1,349 份报告为真实漏洞，其中包括 155 个此前未知的漏洞，其中 108 个已被分配 CVE/PSV 标识符。在 60 个已知漏洞上，Bond 取得了 91.67% 的召回率。与四种领先的 IoT 模糊测试器相比，Bond 将漏洞验证能力最多提升 5.5 倍。消融研究进一步证明了 Bond 关键组件与约束提取的有效性。这些结果确立了 Bond 作为验证固件污点分析结果的实用且有效的框架。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_peng-jiaqian.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_peng-jiaqian.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">2. Heli: Heavy-Light Private Aggregation</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Ryan Lehmkuhl and Henry Corrigan-Gibbs (麻省理工学院); Emma Dauterman (斯坦福大学); David J. Wu (德克萨斯大学奥斯汀分校)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出 Heli，一种允许一对服务器收集关于客户端所持有私密数据的聚合统计、却不获取任何单个客户端数据更多信息的系统。与已有系统一样，Heli 在面对恶意服务器时保护客户端隐私，在面对行为不端的客户端时保护正确性，并支持常见统计函数：均值、方差等。Heli 的创新之处在于，仅其中一个服务器（“重型服务器”）需要执行与客户端数量成正比的每次运行工作；另一个服务器（“轻型服务器”）在一次性设置阶段之后所做的工作与客户端数量无关。因此，一个计算能力受限的参与方，例如预算有限的非营利组织，有可能作为拥有数百万客户端的 Heli 部署中的第二个服务器。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Heli 依赖一种新的密码学原语——仅聚合加密（aggregation-only encryption），允许在许多客户端的加密数据上计算某些受限函数。在拥有一千万客户端的部署中，服务器私下计算 32 个客户端持有的 1 比特整数之和，Heli 的重型服务器完成 240,000 核秒的工作，而轻型服务器完成 7 核毫秒的工作。与已有工作相比，重型服务器的计算量多 38 倍，但轻型服务器的计算量少 120,000 倍。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lehmkuhl.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lehmkuhl.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">3. &#34;Your imaging may be stone-cold normal, but if they look sick, they&#39;re going to get admitted&#34;: An Investigation of Clinicians&#39; Perceptions of Impact &amp; Likelihood of Security Failures</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Ronald E. Thompson III and Hamza Khalid (塔夫茨大学); Hilary Fisher (布里格姆妇女医院); Rhea Votipka (贝斯以色列莱伊健康); Daniel Votipka (塔夫茨大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">网络攻击是关键的病人安全问题，然而安全控制措施往往未能考虑到临床环境的独特性。本文通过混合方法研究来填补对临床医生安全认知的理解空白，首先对美国临床医生进行了 12 次访谈，随后在美国、英国和加拿大开展了一项有 303 名参与者参与的临床医生调查。我们的发现揭示了所感知的威胁与已部署控制措施之间存在显著错位。临床医生认为机密性失效（例如数据泄露）最可能发生。他们将完整性失效（例如被篡改的数值）视为灾难性事件，但相信凭自身专业知识可以忽略异常数据。最后，他们用纸质记录等模拟替代方案来应对既可能又危险的可用性失效，却引入了新的风险。这些结果表明需要将临床医生纳入安全体系，指出现有方法的不足，并为开发更有效的、以临床医生为中心的安全措施提供建议。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_thompson-iii.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_thompson-iii.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">4. Hop: A Modern Transport and Remote Access Protocol</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Paul Flammarion (斯坦福大学); George Hosono (佐治亚理工学院); Wilson Nguyen, Laura Bauman, Daniel Rebelsky, and Gerry Wan (斯坦福大学); David Adrian (独立研究者); Zakir Durumeric (斯坦福大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">自 SSH 标准化以来近 20 年间，对远程访问协议的现实需求以及我们对如何构建安全密码学网络协议的理解都已发生显著演进。在本工作中，我们引入 Hop，一种旨在满足当今需求的传输与远程访问协议。基于现代密码学进展，Hop 降低了 SSH 协议的复杂性与开销，同时通过密码学中介的委托机制、汲取 TLS 与 ACME 经验的原生主机标识、面向现代企业环境的客户端认证，以及对客户端漫游与间歇性连接的支持，解决了 SSH 的诸多不足。我们提出现代远程访问协议的具体设计需求，描述所提出的协议，并评估其性能。我们希望本工作能促进关于未来现代远程访问协议应有形态的讨论。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_flammarion.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_flammarion.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">5. Interpolation-Based Optimization for Enforcing lp-Norm Metric Differential Privacy in Continuous and Fine-Grained Domains</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chenxi Qiu (北德克萨斯大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">度量差分隐私（mDP）通过基于成对距离调整隐私保证来推广局部差分隐私（LDP），从而实现情境感知保护并提升效用。虽然已有的基于优化的方法在粗粒度域中能有效减少效用损失，但由于构造稠密扰动矩阵与满足逐点约束的计算代价，在细粒度或连续场景下优化 mDP 仍具挑战性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出一种基于插值的框架，用于在此类域中优化 ℓp 范数 mDP。我们的方法在一组稀疏的锚点处优化扰动分布，通过对数凸组合在非锚点位置插值分布，可证明地保持 mDP。为解决高维空间中朴素插值导致的隐私违规，我们将插值过程分解为一系列一维步骤，并推导出一个修正公式，从设计上强制满足 ℓp 范数 mDP。我们进一步探讨了跨维度上扰动分布与隐私预算分配的联合优化。在真实位置数据集上的实验表明，我们的方法在细粒度域中提供了严格的隐私保证和有竞争力的效用，优于基线机制。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qiu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qiu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">6. Membership Inference Attacks on Tokenizers of Large Language Models</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Meng Tong (中国科学技术大学); Yuntao Du (普渡大学); Kejiang Chen and Weiming Zhang (中国科学技术大学); Ninghui Li (普渡大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">成员推理攻击（MIA）被广泛用于评估与机器学习模型相关的隐私风险。然而，当这些攻击应用于预训练大语言模型（LLM）时，会遭遇重大挑战，包括样本标注错误、分布偏移以及实验与真实场景下模型规模的差异。为克服这些局限，我们引入分词器（tokenizer）作为成员推理的新型攻击向量。具体而言，分词器将原始文本转换为 LLM 所用的词元。与完整模型不同，分词器可以从头高效训练，从而避免上述挑战。此外，分词器的训练数据通常能代表用于预训练 LLM 的数据。尽管有这些优势，分词器作为攻击向量的潜力仍未被探索。为此，我们首次研究了通过分词器的成员泄露，并探索了五种攻击方法来推断数据集成员关系。在数百万互联网样本上的大量实验揭示了最先进 LLM 分词器中的漏洞。为缓解这一新兴风险，我们进一步提出一种自适应防御方法。我们的发现强调分词器是一种被忽视却至关重要的隐私威胁，凸显了专门为其设计隐私保护机制的迫切需求。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_tong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_tong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">7. JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit Analyses</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu, Wenhui Zhang, Qinglong Wang, and Rui Zheng (浙江大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）展现出令人瞩目的性能，但仍易受越狱攻击，即通过精心构造的对抗性提示绕过安全对齐并诱发出预期之外的响应。尽管越狱攻击普遍存在，其背后的机制仍不甚明了。近期研究主要关注静态表示偏移或识别与生成安全相关的组件。然而，这些研究既未探讨多样的越狱模式，也未提供从电路失效到表示变化的细粒度解释，在揭示越狱机制方面留下了重大空白。在本文中，我们提出 JailbreakScope，一个从表示（越狱如何扭曲 LLM 的危害感知）和电路（越狱如何影响对生成安全至关重要的电路）两个视角分析越狱机制的解释框架，并追踪其在整个生成过程中的演变。我们在 5 个主流 LLM 上、7 种越狱策略下进行了深入评估。我们的评估揭示了一个普遍模式：越狱放大了强化肯定响应的组件，同时抑制了产生拒绝的组件，这使表示向安全区域偏移，导致 LLM 提供响应而非拒绝。此外，我们发现在多样的越狱和多个 LLM 中，表示欺骗与电路激活偏移之间存在强烈且一致的相关性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_he.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_he.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">8. Imitative Membership Inference Attack</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yuntao Du and Yuetian Chen (普渡大学); Hanshen Xiao (普渡大学 &amp; 英伟达研究院); Bruno Ribeiro and Ninghui Li (普渡大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">成员推理攻击（MIA）通过判定特定查询实例是否属于训练集，来评估目标机器学习模型对其训练数据的泄露程度。最先进的 MIA 依赖于训练数百个与目标模型独立的影子模型，导致显著的计算开销。在本文中，我们引入模仿式成员推理攻击（IMIA），它采用一种新颖的模仿训练技术，策略性地构造少量目标感知的模仿模型，使其紧密复现目标模型的行为以用于推理。大量实验结果表明，IMIA 在多种攻击设置下均显著优于现有 MIA，同时仅需不到最先进方法 5% 的计算开销。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_du.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_du.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">9. SMASH: Scalable Maliciously Secure Hybrid Multi-party Computation Framework for Privacy-Preserving Large Language Models</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yunlv Lv and Rui Zhang (中国科学院信息工程研究所; 网络空间安全防御国家重点实验室; 中国科学院大学网络空间安全学院); Zhiyuan Zhang (马克斯·普朗克安全与隐私研究所); Ziyi Wan (中国科学院信息工程研究所; 网络空间安全防御国家重点实验室; 中国科学院大学网络空间安全学院); Lanxue Zhang (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院); Minhui Xue (澳大利亚联邦科学与工业研究组织 Data61 和负责任人工智能研究（RAIR）中心，阿德莱德大学); Jiangtao Li (华东师范大学); Yanan Cao (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）的迅猛崛起引发了对隐私保护推理的迫切需求。然而，现有的恶意安全多方计算（MPC）框架在扩展到大型模型时面临“性能崩溃”，主要原因是非线性算子的二次（O(n²)）通信开销以及昂贵的秘密共享转换。本文提出 SMASH，一种高可扩展的恶意安全混合 MPC 框架，打破了这些瓶颈。SMASH 引入了一种基于 DFT 的旋转技术和一种轻量级知识零知识证明（ZKPoK）构造来求值非线性运算。该方法首次实现了相对于参与方数量的线性通信复杂度（O(n)），且与函数复杂度无关。此外，SMASH 提供了一套高效转换协议（A2L/L2A 以及基于 SM-LUT 的 A2B/B2A），在不依赖昂贵密码学原语的情况下连接算术域与布尔域。大量基准测试表明，SMASH 在运行时间上比最先进框架（如 MP-SPDZ、MD-ML）最多快 18.9 倍，并实现了最多 103 倍的通信缩减。凭借其常数轮在线阶段和对广域网的低敏感性，SMASH 为安全的地域分布式 LLM 部署铺平了道路，在对抗鲁棒性与实用效率之间实现了前所未有的平衡。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lv.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lv.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">10. XGuardian: Towards Generalized, Explainable and More Effective Server-side Anti-cheat in First-Person Shooter Games</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jiayi Zhang, Chenxin Sun, and Chenxiong Qian (香港大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">瞄准辅助外挂是第一人称射击（FPS）游戏中最普遍且臭名昭著的作弊形式，它帮助作弊者非法暴露对手位置并自动瞄准射击，从而对游戏产业构成重大威胁。尽管已投入大量研究努力来自动检测瞄准辅助外挂，现有工作仍存在框架不可靠、泛化能力有限、开销高、检测性能低以及检测结果缺乏可解释性等问题。在本文中，我们提出 XGuardian，一种服务端通用的、可解释的瞄准辅助外挂检测系统，以克服上述局限。它仅需俯仰角和偏航角两种原始数据输入——这是所有 FPS 游戏必备的数据——来构造新颖的时序特征并描述瞄准轨迹，这对于区分作弊者与正常玩家至关重要。XGuardian 以最新主流 FPS 游戏 CS2 进行评估，并用两款不同游戏验证其泛化能力。在不同游戏上、基于真实与大规模数据集，与已有工作相比，它实现了高检测性能和低开销，展现了广泛的泛化能力和高效性。它能够论证其预测结果，从而缩短人工审核的延迟。我们公开了 XGuardian 及其数据集。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-jiayi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-jiayi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">11. The Prompt Stealing Fallacy: Rethinking Metrics, Attacks, and Defenses</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zehang Deng (斯威本科技大学 和 澳大利亚联邦科学与工业研究组织 Data61); Haoyang Li (香港理工大学); Wanlun Ma (斯威本科技大学); Ruoxi Sun and Derui Wang (澳大利亚联邦科学与工业研究组织 Data61); Minhui Xue (澳大利亚联邦科学与工业研究组织 Data61 和负责任人工智能研究（RAIR）中心，阿德莱德大学); Haibo Hu (香港理工大学); Sheng Wen and Yang Xiang (斯威本科技大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">文生图（T2I）模型日益嵌入创意工作流中，精心构造的提示词作为有价值的知识产权（IP）形式存在。然而，这些模型容易受到提示词窃取攻击（PSA），攻击者旨在重建用于生成图像的原始提示词。在本文中，1）我们指出了当前评估实践中的关键不足，并提出两种改进的度量指标：风格相似度（SS）和一种新颖的提示词显著性（PS）评分，二者共同提供对 PSA 有效性的更忠实评估。现有指标仅依赖文本或图像模态之间原始信息与被盗信息的语义相似性，而新指标 PS 和 SS 以更实用的视角评估攻击有效性，明确考虑修饰词的重要性以及由被盗提示词生成图像的风格复制程度。2）通过使用这些度量进行广泛评估，我们发现现有的 PSA 方法——从白盒设置下的软提示词窃取到黑盒设置下的硬提示词窃取——并不如报道的那样有效，尤其是在恢复高贡献提示词组件方面。我们将其归因于根本性约束：白盒方法存在优化目标不匹配的问题，与词元级视觉语义对齐不佳；而黑盒方法由于与目标 T2I 模型生成过程解耦，经历了严重的信息损失。3）我们进一步引入 PromptThief，一种黑盒 PSA 框架，通过利用以 STS 和 SS 为指导的强化学习来引导高词元级贡献恢复，从而解决先前方法中的信息损失问题。PromptThief 在多项指标和真实场景中均显著优于现有基线。4）我们提出并评估了两种防御机制：一种基于对抗样本的主动方法，以及一种通过特征级提示词水印实现的被动方案。我们的评估表明，主动防御对自适应 PSA 仅提供有限的鲁棒性，凸显了在此方向进一步探索的需求。相比之下，被动水印方案在各种图像变换下均展现出强健且一致的检测性能，为提示词 IP 保护提供了一条实用且可靠的路径。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_deng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_deng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">12. When AIOps Become &#34;AI Oops&#34;: Subverting LLM-driven IT Operations via Telemetry Manipulation</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Dario Pasquini (RSAC 实验室); Evgenios M. Kornaropoulos and Giuseppe Ateniese (乔治梅森大学); Omer Akgul, Athanasios Theocharis, and Petros Efstathopoulos (RSAC 实验室)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">面向 IT 运维的人工智能（AIOps）正通过自动化异常检测、事件诊断和修复来改变组织管理复杂软件系统的方式。现代 AIOps 解决方案日益依赖自主的基于 LLM 的智能体来解读遥测数据并在最少人工干预下采取纠正措施，承诺更快的响应速度和运营成本节约。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们对 AIOps 解决方案进行了首次安全分析，表明 AI 驱动的自动化再次带来深远的安全代价。我们证明，攻击者可以操纵系统遥测数据，误导 AIOps 智能体采取危害其所管理基础设施完整性的行动。我们引入了可靠注入遥测数据的技术，利用诱发错误的请求，通过一种对抗性奖励劫持的形式影响智能体行为；以及看似合理但不正确的系统错误解读来引导智能体的决策。我们的攻击方法 AIOpsDoom 完全自动化——结合侦察、模糊测试和 LLM 驱动的对抗性输入生成——且在毫无目标系统先验知识的情况下运行。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为应对这一威胁，我们提出 AIOpsShield，一种利用遥测数据的结构化特性和用户生成内容最小作用来净化遥测数据的防御机制。我们的实验表明，AIOpsShield 能可靠地阻断基于遥测的攻击，同时不影响智能体的正常性能。最终，本工作揭示了 AIOps 作为系统攻陷新兴攻击向量的风险，强调了安全感知的 AIOps 设计的迫切需求。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_pasquini.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_pasquini.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">13. A Large-Scale Study of Personalized Phishing using Large Language Models</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Stefan Czybik (BIFOLD 和 柏林工业大学); Anne Josiane Kouam (法国国家信息与自动化研究所 和 柏林工业大学); Peter Heubl and Jan Magnus Nold (波鸿鲁尔大学); Konrad Rieck (BIFOLD 和 柏林工业大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）能够生成流畅且具有说服力的文本，使其成为有价值的沟通工具。然而，这一能力也使其对恶意用途具有吸引力。虽然多项研究表明 LLM 可支持通用钓鱼攻击，但其在大规模个性化攻击中的潜力尚未被探索和量化。因此，在本研究中，我们在一项有 7700 名参与者参与的实验中评估了基于 LLM 的鱼叉式钓鱼的有效性。我们以目标电子邮件地址作为查询，通过网络搜索收集个人信息，并自动生成针对每位参与者量身定制的邮件。我们的发现揭示了一个令人担忧的局面：与通用钓鱼策略相比，基于 LLM 的鱼叉式钓鱼使点击率提高了近三倍。这一效应是一致的，无论通用邮件是由人工撰写还是同样由 LLM 生成。此外，个性化的成本极低，每封邮件约 0.03 美元。鉴于钓鱼仍然是针对 IT 基础设施的主要攻击向量，我们得出结论：迫切需要加强现有防御，例如限制与电子邮件地址可关联的公开信息，并将个性化钓鱼纳入安全意识培训。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_czybik.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_czybik.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">14. Secure Protocol Composition under Dynamic Corruption: Scaling Up Symbolic Analysis for Real-World Security Properties</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Cas Cremers, Erik Pallas, and Aleksi Peltonen (德国亥姆霍兹信息安全研究中心（CISPA）)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">尽管自动化符号协议验证已被证明既有价值又有效，但当前方法开始达到其极限：小型协议可以自动分析，但最复杂的案例研究往往需要大量的专家时间和资源。已有许多尝试通过组合验证来解决此问题，但它们依赖于不切实际的协议假设，且不支持前向保密等真实世界安全属性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们使针对真实世界安全协议的现代安全属性的组合符号分析成为可能。我们在 Applied π-Calculus 中发展了一个组合结果，当协议满足不相交性要求时，即使在存在能够动态腐败的攻击者的情况下该结果也成立。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们通过将一个数据交换协议与 Diffie-Hellman 密钥交换进行组合，以及在 RFC 8446 范围和 ECH 扩展内对 TLS 1.3 的前向保密进行组合分析，展示了我们结果的适用性和有效性。虽然对带 ECH 的 TLS 1.3 的单体分析在 10% 的情况下无法给出结果，但所有组合分析均成功。此外，运行时间平均减少 71%，内存使用平均减少 86%。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cremers.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cremers.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">15. Opossum Attack: Application Layer Desynchronization using Opportunistic TLS</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Robert Merget (技术创新研究所); Nurullah Erinola and Marcel Maehren (波鸿鲁尔大学); Lukas Knittel (波鸿鲁尔大学); Sven Hebrok (帕德博恩大学); Marcus Brinkmann (波鸿鲁尔大学); Juraj Somorovsky (帕德博恩大学); Jörg Schwenk (波鸿鲁尔大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">许多协议（如 HTTP、FTP、POP3 和 SMTP）最初被设计为同步明文协议——命令和数据以明文发送，客户端在发送下一个请求之前等待当前待处理请求的响应。后来，引入了两种主要方案来为这些协议加装 TLS 保护。（1）隐式 TLS：为每个基于 TLS 的协议指定一个新的知名 TCP 端口，并立即以 TLS 开始。（2）机会式 TLS：保留原有的知名端口并以明文协议开始，然后响应 STARTTLS 之类的命令切换到 TLS。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们揭示了通过隐式和机会式 TLS 将 TLS 集成到流行应用层协议中的一种新颖弱点。该弱点破坏了认证，即使在现代 TLS 实现中，如果同时支持隐式 TLS 和机会式 TLS 也会受到影响。然后，该认证缺陷可被利用，从纯粹的中间人位置影响 TLS 握手后交换的消息。与先前对机会式 TLS 的攻击不同，此类攻击不依赖于实现中的漏洞，仅需通信一方支持机会式 TLS。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们分析了支持机会式 TLS 的流行应用层协议对该攻击的脆弱性。为展示攻击的实际影响，我们详细分析了 HTTP（RFC 2817）的利用技术，并展示了四个不同的利用方向。为评估攻击对已部署服务器的影响，我们对多种协议和端口进行了一系列 IPv4 全范围扫描，以检查机会式 TLS 的支持情况。我们发现，许多应用协议对机会式 TLS 的支持仍然广泛，超过 300 万台服务器同时支持隐式和机会式 TLS。在 HTTP 方面，我们发现 35 个端口上有 20,121 台服务器支持机会式 HTTP，其中 2,268 台同时支持 HTTPS，539 台使用与隐式 HTTPS 相同的域名，构成可利用场景。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_merget.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_merget.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">16. Breaking Widely Deployed Perceptual Hash Functions: Black-Box Collisions in Apple NeuralHash and Microsoft PhotoDNA</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Diane Leblanc-Albarel and Bart Preneel (鲁汶大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">感知哈希函数被设计用于检测多媒体版权违规和非法内容。为实现其目的，它们将感知上相似的输入映射到相近的输出。然而，对于许多广泛部署的方案，其设计策略和详细规格仍为专有。各国政府现正考虑将其扩展到端到端加密服务的客户端扫描（CSS），在加密前验证内容是否为非法材料。2021 年，Apple 提出了一项基于 NeuralHash 感知哈希函数的 CSS 详细提案。在隐私和安全方面的强烈批评后，Apple 撤回了该提案，但 NeuralHash 仍部署在所有设备上，其当前用途未予披露。理论上，对 NeuralHash（96 位哈希值）的暴力碰撞需要 2^48 次求值。NeuralHash 发布后不久，研究人员表明很容易构造感知上不相似的碰撞，通过发送与非法内容共享相同哈希值的无辜图像来诬陷任何用户。本工作揭示了一个更严重的弱点：当输入限定为人脸时，我们仅在 2^16 次哈希函数求值后就发现了多对感知上不同图像之间的碰撞。与定向攻击不同，我们的黑盒方法无需了解哈希函数设计。我们还展示了高假阴性率（本应共享相同哈希却不共享的图像）。我们通过研究 Microsoft 广泛部署的 1152 位感知哈希函数 PhotoDNA 进一步证实了该方法的普适性。在 PhotoDNA 方面，我们在远低于此前报道的阈值下发现了近碰撞，根据所用阈值不同，出现在 2^14.6 到 2^17 次求值之间。这是首项在 NeuralHash 中展示精确碰撞、并在如此低阈值下识别 PhotoDNA 近碰撞的工作。这些结果对这些设计是否适用于大规模客户端扫描提出了严重质疑，因为它们产生高假阳性率和假阴性率，并凸显了重新评估其安全性和可行性的必要性，尤其是在隐私风险和假阳性具有严重后果的大规模应用中。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_leblanc-albarel.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_leblanc-albarel.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">17. Zero Knowledge (About) Encryption: A Comparative Security Analysis of Three Cloud-based Password Managers</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Matteo Scarlata (苏黎世联邦理工学院); Giovanni Torrisi and Matilda Backendal (瑞士意大利语区大学，卢加诺); Kenneth G. Paterson (苏黎世联邦理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">零知识加密是云密码管理器厂商广泛使用的一个术语。虽然该术语没有严格的技术含义，但它传达了这样的理念：代表用户存储加密密码保险库的服务器无法获知这些保险库的内容。厂商所做的安全声明意味着，即使服务器完全是恶意的，这一点也应成立。这一威胁模型在实践中是合理的，因为保险库数据的高敏感性使密码管理器服务器成为攻击的诱人目标（攻击历史也证明了这一点）。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们考察了针对完全恶意服务器的安全保护在多大程度上成立，针对的是三家做出零知识加密声明的领先厂商：Bitwarden、LastPass 和 Dashlane。它们合计拥有超过 6000 万用户和 23% 的市场份额。我们提出了针对 Bitwarden 的 12 种不同攻击、针对 LastPass 的 7 种攻击和针对 Dashlane 的 6 种攻击。这些攻击的严重程度各异，从对目标用户保险库的完整性违规到对与某组织关联的所有保险库的完全攻陷。大多数攻击允许恢复密码。我们已向厂商披露了发现，修复工作正在进行中。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的攻击展示了为云密码管理器考虑恶意服务器威胁模型的重要性。尽管厂商努力在此场景下实现安全性，我们仍发现了若干导致漏洞的常见设计反模式和密码学误解。我们讨论了可能的缓解措施，并更广泛地反思了端到端加密系统开发者可从我们的分析中汲取的教训。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_scarlata.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_scarlata.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">18. A Midsummer Meme&#39;s Dream: Investigating Market Manipulations in the Meme Coin Ecosystem</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Alberto Maria Mongardini (罗马大学和丹麦技术大学); Alessandro Mei (罗马大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">从病毒式笑话到价值数十亿美元的现象，模因币已成为加密货币市场中最受欢迎的板块之一。与比特币等注重实用性的加密资产不同，模因币的价值主要源自社区情绪，使其容易受到操纵。本研究对模因币生态系统进行了前所未有的跨链分析，考察了 Ethereum、BNB Smart Chain、Solana 和 Base 上共计 34,988 个代币。我们刻画了它们的代币经济学特征，并在为期三个月的纵向分析中追踪其增长态势。我们发现，在高回报代币（&gt;100%）中，高达 82.89% 存在人为增长策略的迹象，这些策略旨在制造市场关注度的误导性假象。其中包括洗售交易以及一种我们定义的新型操纵方式——基于流动性池的价格膨胀（LPI），即通过少量战略性买入引发价格剧烈上涨。我们发现，利润攫取手段（如拉高出货和跑路）通常紧随洗售交易或 LPI 等初始操纵行为之后，这表明早期操纵为后续的利用与收割奠定了基础。我们对这些手段的经济影响进行了量化，识别出超过 17,000 个受害者地址，其已实现损失超过 930 万美元。这些发现揭示了组合式操纵在高表现模因币中广泛存在，表明其惊人的涨幅往往源于协调性行动，而非自然的市场动态。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_mongardini.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_mongardini.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">19. Distributed Vector Commitments and Their Applications</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Rui Gao and Huaqun Wang (南京邮电大学，西藏智能国家重点实验室); Zhiguo Wan (杭州师范大学); Yuncong Hu (上海交通大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">向量承诺（VC）方案允许证明者承诺一个向量，并在随后以简短证明打开任意位置。然而，现有的 VC 方案是为中心化场景设计的，无法用于输入向量分布在多台机器上的去中心化系统。同样，传统 VC 方案也无法利用跨多台机器的分布式并行计算来加速。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为解决这一问题，我们引入了一个新概念——分布式 VC（DVC），它允许多台机器（每台仅持有输入向量的一个子向量）协同承诺整个向量，并以分布式方式生成位置证明。据我们所知，目前尚无关于 DVC 的先前工作，也没有现有工作可以轻易推导出高效的 DVC 方案。其关键挑战在于，承诺和证明均依赖于整个向量，而在分布式场景中没有任何单台机器持有完整的向量。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了首个 DVC 方案 HLE-DVC，它利用 M 台机器并行处理长度为 N 的分布式向量 v，每台机器持有长度为 N/M 的子向量。HLE-DVC 实现了紧凑的证明大小 O(log M)，并允许每台机器在单轮通信中生成其所有位置证明，通信开销为 O(log M)，计算开销为 O(N log N / M)。此外，HLE-DVC 支持批量证明、证明聚合和高效更新。我们进行了实验并开源了代码。使用 256 台机器为长度为 2^30 的承诺向量生成所有证明耗时 17,515 秒。这相比单机上的 HLE-DVC 实现了 256 倍的并行加速，比 Hyperproofs（著名的单机 VC 方案）快 142 倍。每台机器的通信开销为 0.768 KB。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">20. From Mirai to Gorilla: Deep Dive into a Long-Lasting DDoS-for-Hire Botnet</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Maarten Weyns, Dario Ferrero, and Stefan Op de Beek (代尔夫特理工大学); Daniel Wagner (马克斯·普朗克信息学研究所 / DE-CIX); Georgios Smaragdakis and Harm Griffioen (代尔夫特理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">2016 年，Mirai 僵尸网络席卷互联网，开启了 DDoS 攻击的新时代。在随后的十年中，Mirai 僵尸网络的衍生品从简单的攻击工具转变为商业化平台，提供分布式拒绝服务攻击即服务。这类平台使用户能够以极低的技术门槛发起大规模 DDoS 攻击。一个典型的例子是 Gorilla 僵尸网络，其运营时间从 2024 年秋季持续至 2025 年夏季，与同类基于 Mirai 的僵尸网络相比，其生命周期异常漫长。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们对基于 Mirai 的 Gorilla 僵尸网络进行逆向工程，旨在理解其设计、工程决策和营销策略，以揭示其为何具有如此强的韧性和成功率。我们调查了其运营特征，包括所支持的攻击类型、底层基础设施以及僵尸节点的行为。我们发现，Gorilla 的长寿命源于有针对性的改进，包括两个软件开发阶段以及对先前版本经验的学习，这使其区别于典型的 Mirai 僵尸网络。在此过程中，我们分析了 Gorilla 僵尸网络的攻击火力和攻击向量，并刻画了其目标的业务类型。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_weyns.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_weyns.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">21. LPG: Raise Your Location Privacy Game in Direct-to-Cell LEO Satellite Networks</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Quan Shi (新加坡国立大学); Liying Wang (北京大学); Prosanta Gope (谢菲尔德大学); Qi Liang and Haowen Wang (北京邮电大学); Qirui Liu and Chenren Xu (北京大学); Shangguang Wang and Qing Li (北京邮电大学); Biplab Sikdar (新加坡国立大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">多租户直连蜂窝（D2C）低轨道（LEO）卫星网络通过将移动网络运营商（MNO）管理的身份与卫星网络运营商（SNO）可见的位置相关联，对用户的位置隐私构成了重大风险。现有的隐私保护方案不适用于这些卫星环境中资源受限的硬件和轨道动态特性。我们提出了 LPG（Location Privacy Game），这是首个为 D2C LEO 提供用户可配置位置隐私的协议层解决方案。LPG 通过身份-位置解耦来实现这一目标：SNO 在不可见用户身份的情况下提供连接，而 MNO 在无法获取精确位置信息的情况下管理服务和计费。LPG 支持离线安全认证和密钥协商而无需向卫星泄露用户身份，支持按所选地理粒度进行用户可配置的位置披露以满足基本服务需求，并通过隐私保护的结算方式确保 MNO 与 SNO 之间的公平计费。我们在真实在轨 LEO 卫星和商用手机上的实现表明，LPG 在资源受限、高度动态的 LEO 环境中是实用且可行的。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">22. PrivacyShield: Relaying BLE Beacons to Counter Unsolicited Tracking</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Florian Hofhammer (洛桑联邦理工学院); Daniele Antonioli (EURECOM); Mathias Payer (洛桑联邦理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">离线查找网络（如 Apple 的 Find My、Google 的 Find My Device 或 Samsung 的 SmartThings Find）经常被滥用于跟踪毫无防备的受害者。这些网络允许用户将小巧廉价的标签附着在物品上，以便在物品丢失时定位。标签通过蓝牙低功耗（BLE）信标广播其存在，附近联网的设备（如智能手机）将其位置报告给查找网络。然而，离线查找标签的低廉价格和易于隐藏的特性使其对恶意行为者具有吸引力，他们会将标签放置在不知情的受害者身上。附近的设备甚至受害者自己的设备随后会在不知不觉中将受害者的位置报告给跟踪者。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们分析了离线查找网络所采取的防跟踪措施，重点关注 Apple 的 Find My 和 Google 的 Find My Device。我们展示了恶意行为者如何绕过这些措施，并提出了 PrivacyShield，一种保护跟踪受害者的新型中继网络。我们的网络利用了离线查找 BLE 信标未经认证且可被中继到任意位置这一事实。被中继的信标会致使第三方设备向查找网络报告错误的位置，从而混淆受害者的真实位置。我们演示了 PrivacyShield 在隐藏标签位置方面的有效性，并展示了系统在面对阻止其使用的尝试时的鲁棒性。随后，我们为离线查找网络提供商提出了改进跟踪保护的实际建议。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hofhammer.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hofhammer.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">23. Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Vijay Prakash (纽约大学); Majed Almansoori (威斯康星大学麦迪逊分校); Donghan Hu (纽约大学); Rahul Chatterjee (威斯康星大学麦迪逊分校); Danny Yuxing Huang (纽约大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">技术介导的虐待（TFA）是亲密伴侣暴力（IPV）的一种普遍形式，利用数字工具来控制、监视或伤害幸存者。虽然科技诊所是 TFA 幸存者可靠的支持来源之一，但由于人员配置限制和后勤障碍，它们面临诸多局限。因此，许多幸存者转而寻求在线资源的帮助。随着大型语言模型（LLM）的可及性和普及度不断提高，以及 IPV 组织日益增长的兴趣，幸存者可能会在寻求科技诊所帮助之前先咨询基于 LLM 的聊天机器人。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在这项工作中，我们首次对四个 LLM 进行了由专家主导的人工评估——包括两个广泛使用的通用非推理模型和两个专为 IPV 场景设计的领域专用模型——重点关注其在回答 TFA 相关问题方面的有效性。我们使用从文献和在线论坛收集的真实问题，在为 TFA 领域量身定制的评估标准下，评估了以幸存者安全为中心的提示所生成的零样本单轮 LLM 响应的质量。此外，我们开展了一项用户研究，从经历过 TFA 的个体视角评估这些响应的可感知可操作性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的发现基于专家评估和用户反馈，为 LLM 在 TFA 场景中的当前能力和局限性提供了深入见解，可为该领域未来模型的设计、开发和微调提供参考。最后，我们提出了改进 LLM 在幸存者支持方面表现的具体建议。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_prakash.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_prakash.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">24. Residual-PAC Privacy: Automatic Privacy Control Beyond the Gaussian Barrier</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Tao Zhang and Yevgeniy Vorobeychik (圣路易斯华盛顿大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">概率近似正确（PAC）隐私框架 [xiao2023pac] 提供了一种强大的基于实例的方法论，用于在复杂的数据驱动系统中保护隐私。现有的 PAC 隐私算法（我们称之为 Auto-PAC）依赖于高斯互信息上界。然而，我们证明 Auto-PAC 所获得的上界当且仅当在数据分布下未扰动的输出为高斯分布且噪声为独立高斯分布时才是紧的。我们提出了两种解决该问题的方法。首先，我们引入了两种可处理的 Auto-PAC 后处理方法，分别基于 Donsker–Varadhan 表示和切片 Wasserstein 距离。然而，该结果仍然存在“浪费”的隐私预算。为了更根本地解决这一问题，我们引入了 Residual-PAC（R-PAC）隐私，一种基于 f-散度的度量，用于量化对抗推理之后剩余的隐私。为了在实践中实现 R-PAC 隐私，我们提出了 Stackelberg Residual-PAC（SR-PAC）自动隐私化算法，这是一个博弈论框架，通过凸双层优化选择最优的噪声分布。我们的方法对任意数据分布都能实现高效的隐私预算利用，并在多个机制访问同一数据集时自然地组合。我们的实验表明，SR-PAC 始终比 PAC 和差分隐私基线取得更好的隐私-效用权衡。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-tao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-tao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">25. HAMLOCK: HArdware-Model LOgically Combined attacK</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Sanskar Amgain (田纳西大学诺克斯维尔分校); Daniel Lobo, Atri Chatterjee, and Swarup Bhunia (佛罗里达大学); Fnu Suya (田纳西大学诺克斯维尔分校)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">第三方硬件加速器（如 FPGA、ASIC）在深度神经网络（DNN）中日益广泛的使用带来了新的安全漏洞。当前的模型级后门攻击仅通过投毒模型权重来对带有特定触发器的输入进行错误分类，这种做法将完整的逐层后门激活路径嵌入模型内部，因此往往能被最先进的防御方法检测到。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了硬件-模型逻辑组合攻击（HAMLOCK），一种更为隐蔽的威胁，将攻击逻辑分散在硬件-软件边界上。软件（模型）仅通过微调少量神经元的激活值进行最小程度的修改，使其在触发器存在时产生独特的高激活值。恶意硬件木马通过监测相应神经元的最高有效位或 8 位指数来检测这些独特激活，并触发另一个硬件木马直接操纵最终输出 logits 以实现错误分类。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">这种解耦设计具有极高的隐蔽性，因为模型本身不像常规攻击那样包含完整的后门激活路径，因而看起来完全良性。在实验中，在 MNIST、CIFAR10、GTSRB 和 ImageNet 等基准测试上，HAMLOCK 实现了近乎完美的攻击成功率，且干净准确率下降微乎其微。更重要的是，HAMLOCK 无需任何自适应优化即可绕过最先进的模型级防御。该硬件木马同样无法被检测，其面积和功耗开销低至 0.01%，极易被工艺和环境噪声所掩盖。我们的发现揭示了硬件-软件接口处的一个关键漏洞，亟需针对这一新兴威胁开发新的跨层防御。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_amgain.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_amgain.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">26. Vεrity: Verifiable Local Differential Privacy</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>James Bell-Clark, Adrià Gascón, Baiyu Li, and Mariana Raykova (谷歌); Amrita Roy Chowdhury (密歇根大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本地差分隐私（LDP）使个体能够在保护隐私的前提下报告敏感数据。然而，LDP 机制容易受到投毒攻击的影响：控制部分报告用户的对抗者可以显著扭曲聚合输出——其扭曲程度远甚于输入直接报告的非隐私方案。在本文中，我们提出了两种新颖的解决方案，在保持 LDP 隐私保证的同时防止投毒攻击。第一种方案 Vεrity-Auth 针对用户报告的输入存在第三方可用的基本真值的场景。第二种方案 Vεrity 解决更具挑战性的情形：用户在本地生成输入，且不存在可用于引导可验证随机数生成的基本真值。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bell.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bell.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">27. Missing, Present and Conflicting: A Large Scale Analysis of IoT Update Information in the EU Market</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Swaathi Vetrivel, Michel van Eeten, and Carlos H. Gañán (代尔夫特理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">安全更新对于保护 IoT 设备至关重要，然而消费者往往缺乏关于设备将获得多长期限支持的可靠信息。我们对欧洲市场的更新期限披露情况进行了首次大规模研究，分析了本地零售商、欧盟 Amazon 站点和 Temu 上共计 34,187 个产品页面。披露情况差异显著：受监管监督的荷兰零售商高达 92% 的设备列出了更新期限，而 Amazon 提供此类信息的比例不足 1%，Temu 则完全没有。对于欧盟规则强制要求披露的智能电视，覆盖率较高但仍然不一致。所声明的更新期限从一年到八年不等，其中智能电视通常获得最长的支持。通过比较各零售商、制造商以及欧盟中央产品数据库所声明的支持期限，我们发现了广泛的矛盾之处，零售商往往相对于制造商低估了支持期限。这些不一致性限制了透明度要求的效力，并可能误导消费者。我们的发现表明，监管可以提高可见性，但只有强有力的执行和标准化的披露机制才能确保信息的准确和可信。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_vetrivel.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_vetrivel.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">28. VSG-Safe: Spotting NSFW Video through Cross-Frame Evidence</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yuyang Zhang, Xudong Jiang, Yuxuan Song, and Yuxiang Sun (武汉大学); Yihao Huang (新加坡国立大学); Run Wang, Shundi Xiao, and Lina Wang (武汉大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">文本到视频（T2V）模型的最新进展使得生成高保真且紧密遵循文本提示的视频成为可能。然而，这在扩展实际应用的同时，也放大了由视觉内容自动合成所带来的严重安全和社会问题——这些内容在某些使用场景（如公共或办公场所）中可能是不当的，包括色情或暴力内容（例如，Grok 可以在“Spicy”模式下生成色情视频）。我们观察到，此类视觉内容通常分布在多帧之间，嵌入在视觉实体、其属性以及实体间关系中。而现有的内容审核流水线主要将视频内容视为独立的单帧或原始帧序列，忽视了关键语义可以通过特定帧的组合来体现这一事实。这一缺陷使得它们无法进行跨帧推理，将检测局限于低层视觉线索（如血腥画面或显式冲突），导致在需要跨帧推理时（包括违法活动或威胁）频繁失效。为解决这些局限性，我们提出利用场景图作为核心的中间语义表示。场景图自然地编码了实体、其属性和实体间关系，同时支持对跨帧内容的推理。基于这一洞察，我们进一步提出了 VSG-Safe，一种面向 T2V 内容审核的新型场景图驱动框架。具体而言，我们的方法首先从视频中提取跨帧内容以构建场景图。借助这些图，我们利用面向图的模型联合捕获实体、属性和实体间关系，从而实现有效检测。为评估其有效性，我们在 SOTA 基准测试和我们自构建的视频数据集上进行了大量实验。VSG-Safe 的平均 F1 分数达到 97.62%，比七个基线方法平均高出 42.32%。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-yuyang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-yuyang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">29. Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining Framework</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Wenbo Guo (南洋理工大学); Chengwei Liu (南开大学); Ming Kang (四川大学); Yiran Zhang and Jiahui Wu (南洋理工大学); Zhengzi Xu (Imperial Global Singapore); Vinay Sachidananda and Yang Liu (南洋理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Python Package Index（PyPI）已成为恶意行为者的攻击目标，然而现有检测工具的误报率高达 15-30%，错误地将三分之一的合法包标记为恶意包。这一问题源于当前工具依赖简单的语法规则而非语义理解，无法区分服务于合法目的还是恶意目的的相同 API 调用。为应对这一挑战，我们提出了 PyGuard，一种知识驱动框架，通过从现有工具的假阳性和假阴性中提取模式，将检测失败转化为有用的行为知识。该方法使用分层模式挖掘来识别区分恶意代码与良性代码的行为序列，利用大型语言模型创建超越语法变体的语义抽象，并将这些知识整合到一个融合精确模式匹配与上下文推理的检测系统中。PyGuard 达到了 99.50% 的准确率，仅有 2 个假阳性，而现有工具为 1,927-2,117 个；在混淆代码上保持 98.28% 的准确率；并在实际部署中识别出 219 个此前未知的恶意包。这些行为模式展现了跨生态系统的适用性，在 NPM 包上达到 98.07% 的准确率，证明了语义理解能够实现跨编程语言的知识迁移。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-wenbo.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-wenbo.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">30. End-to-End Encrypted Collaborative Documents</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Christian Knabenhans (洛桑联邦理工学院); Zayd Maradni (马克斯·普朗克软件系统研究所); Carmela Troncoso (马克斯·普朗克安全与隐私研究所 &amp; 洛桑联邦理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">协作文档（如 Google Docs、Microsoft 365）通常包含敏感信息，如个人或财务数据。在这项工作中，我们将目前（主要）局限于消息传递场景的端到端加密（E2EE）保护扩展到协作文档。我们梳理并形式化了端到端加密协作文档（E2EE-CD）的安全和功能需求。随后，我们提出了一个实现 E2EE-CD 的通用框架，通过将端到端加密异步广播通道与任何确保文档全局一致性视图的编辑协调机制相结合来实现。我们给出了形式化证明，将 E2EE-CD 方案的安全性直接关联到底层端到端加密通信通道的安全性。随后，我们梳理了调查记者使用 E2EE-CD 的额外部署需求，并设计了 SignalCD，这是一种构建于 Signal 群组消息协议之上、针对该场景量身定制的 E2EE-CD 系统。我们分析了 SignalCD 的安全保证，实现了原型系统，并通过实验证明我们的解决方案效率足以支持实时协作。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_knabenhans.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_knabenhans.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">31. SophOMR: Improved Oblivious Message Retrieval from SIMD-Aware Homomorphic Compression</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Keewoo Lee (以太坊基金会); Yongdong Yeo (首尔国立大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">保护隐私的区块链和确保接收者隐私的私密消息服务面临一项重大的用户体验挑战：每个客户端必须扫描公共公告板上的所有载荷（payload），以免遗漏发给自己的消息。不经意消息检索（OMR）通过使用同态加密（HE）将这一昂贵的扫描过程安全地外包给服务提供商来解决该问题。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在这项工作中，我们提出了一种新的 OMR 方案，在先前最先进的方案 PerfOMR（USENIX Security&#39;24）的基础上实现了大幅改进。我们的实现在一个包含 65,536 个载荷（每个 612 字节，其中最多 50 个相关）的场景中，运行时间减少了 3.4 倍，摘要大小减少了 2.2 倍，密钥大小减少了 1.5 倍。这些改进的核心是一种新的同态压缩机制，其中长度与载荷总数成正比的密文被压缩为长度与相关载荷数量上限成正比的摘要。与先前的方法不同，我们的方案充分利用了底层 HE 方案的原生同态 SIMD 结构，显著提升了效率。在上述场景中，我们的压缩方案相比 PerfOMR 实现了 7.5 倍的加速。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lee.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lee.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">32. FABS: Fast Attribute-Based Signatures</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Liqun Chen, Long Meng, Yalan Wang, Nada El Kassem, Christopher JP Newton, Yangguang Tian, Jodie Knapp, Constantin Cătălin Drăgan, and Daniel Gardham (萨里大学); Mark Manulis (慕尼黑联邦国防军大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于属性的签名（ABS）提供了对谁可以生成数字签名的细粒度控制，在许多实际场景中有广泛应用。本文提出了一对快速的 ABS 方案：一个用于密钥策略 ABS（KP-ABS），另一个用于签名策略 ABS（SP-ABS）。两个方案均使用单调跨度程序（MSP）支持表达性策略，并提供了实用特性，如大属性域、任意属性和自适应安全。最值得注意的是，我们提供了基于 MSP 的 ABS 方案的首个实现，并证明我们的方案在该领域实现了已知的最佳渐近性能和具体性能。在渐近性方面，密钥生成、签名和验证时间与属性数量呈线性关系；验证仅需两次配对运算。在具体性能方面，对于 100 个属性，我们的 KP-ABS 方案分别用时 0.16 秒、0.10 秒和 0.13 秒完成密钥生成、签名和验证；我们的 SP-ABS 方案在相同操作上分别用时 0.082 秒、0.26 秒和 0.21 秒。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-liqun.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-liqun.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">33. Invariant-Guided Logical Testing of Open RAN Controllers</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Tianchang Yang, Ali Ranjbar, Gang Tan, and Syed Rafiul Hussain (宾夕法尼亚州立大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Open RAN（O-RAN）代表了移动网络架构的根本性变革，通过开放接口和软件驱动组件推进互操作性和灵活性。这一变革在带来可编程性和创新的同时，也使得 O-RAN 组件的逻辑正确性对于网络的安全可靠运行变得至关重要。然而，由于系统复杂性、实现多样性以及缺乏明确的正确性判定准则，验证 O-RAN 的语义正确性仍然具有挑战性。我们提出了 InvaRAN，一个系统性测试框架，利用动态推断的程序不变量作为预期行为的代理来检测 O-RAN 实现中的逻辑缺陷。为减少误报并聚焦于语义上有意义的行为，InvaRAN 根据不变量对程序逻辑的影响将其分为关键和非关键两类。超越了仅推断有限语义关系的传统基于模板的不变量推断方法，InvaRAN 捕获跨执行踪迹的变量间关联以发现更具表达力的语义联系。我们在两个生产级 O-RAN 控制器的平台组件和 xApps 上评估了 InvaRAN。InvaRAN 发现了九个此前未知的问题，包括七个逻辑漏洞和两个内存漏洞，证明了不变量引导的测试在暴露 O-RAN 系统中隐蔽的、规范未涉及的缺陷方面的有效性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_yang-tianchang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_yang-tianchang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">34. Analyzing the WebRTC Ecosystem and Breaking Authentication in DTLS-SRTP</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Martin Bach (技术创新研究所); Vukašin Karadžić (达姆施塔特工业大学); Lukas Knittel (波鸿鲁尔大学); Robert Merget and Jean Paul Degabriele (技术创新研究所)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">DTLS-SRTP 旨在保护实时媒体通信安全，被广泛应用于 Zoom、Teams 和 Google Meet 等知名音视频通话平台。值得注意的是，它是 Web 实时通信（WebRTC）标准的一部分，该标准使浏览器中的实时通信成为可能。为此，WebRTC 使用了多种技术，包括 HTTP、TLS、SDP、ICE、STUN、TURN、UDP、TCP、DTLS、(S)RTP、(S)RTCP 和 SCTP。这种技术的混合导致系统过于复杂，难以进行系统化和自动化审计。因此，这一核心现代通信技术的部署安全性在很大程度上仍未被探索。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在这项工作中，我们旨在填补这一空白，开发了一个自动化中间人（MitM）测试框架（DTLS-MitM-Scanner（DMS）），用于测试 DTLS-SRTP 连接中 DTLS 通道的安全性。我们使用该框架在一项涵盖 24 家服务提供商的浏览器和移动应用的研究中对生态系统的现状进行了分析。我们的分析特别关注 DTLS-SRTP 中的认证机制，测试了 19 种可能导致客户端或服务器认证绕过的潜在漏洞。我们发现，在 33 个受测试的媒体服务器实现中，有 19 个存在漏洞，允许攻击者在 DTLS 层攻破认证。对于其中 9 个受影响的系统（服务于数亿用户），我们还进一步证明，在仅具备中间人能力的假设下，攻击者可以利用这些漏洞获取媒体数据。我们通过构建一个概念验证漏洞利用来窃听 Webex 视频会议通话，以凸显这些漏洞的影响。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bach.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bach.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">35. Security and Privacy Analysis of Tile&#39;s Location Tracking Protocol</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Akshaya Kumar, Anna Raymaker, and Michael A. Specter (佐治亚理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们对 Tile 进行了首次全面的安全分析。Tile 是仅次于 Apple AirTags 的第二大受欢迎的众包位置追踪服务。我们识别出多个可利用的漏洞和设计缺陷，推翻了该平台声称的诸多安全与隐私保证：Tile 的服务器可以持续获知所有用户和标签的位置，无权限的攻击者可以通过 Tile 设备发出的 Bluetooth 广告追踪用户，而且 Tile 的防盗模式很容易被攻破。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">尽管部署规模庞大——拥有数百万用户、设备以及专用硬件标签——Tile 却未提供对其协议或威胁模型的正式描述。更糟糕的是，Tile 为了支持防盗用例而有意削弱其反跟踪功能，并依赖一种新颖的&#34;问责&#34;机制来惩罚那些滥用系统跟踪受害者的人。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们考察了 Tile 的问责机制，这一独特功能具有独立的研究价值；没有其他提供商尝试保证问责。虽然理想的问责机制可能遏制众包位置追踪协议中的滥用行为，但我们表明 Tile 的实现是可以被攻破的，并引入了新的可利用漏洞。最后，我们讨论了在该场景下建立新的、正式的问责定义的必要性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kumar.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kumar.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">36. VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision–Language Inference</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Minfeng Qi and Dongyang He (澳门城市大学); Qin Wang (澳大利亚联邦科学与工业研究组织 Data61); Lefeng Zhang (澳门城市大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">视觉推理验证码（Visual Reasoning CAPTCHAs，VRC）将视觉场景与要求对对象、属性和空间关系进行组合推理的自然语言查询相结合。它们越来越多地被部署为抵御自动化机器人的主要防御手段。现有的求解器分为两类范式：以视觉为中心的方法依赖特定模板的检测器，但难以应对新布局；以推理为中心的方法利用 LLM，但在细粒度视觉感知方面表现不佳。两者都缺乏处理异构 VRC 部署所需的通用性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出 VIPER，一个将结构化多对象视觉感知与基于 LLM 的自适应推理相结合的统一攻击框架。VIPER 解析视觉布局，将属性与问题语义关联，并在模块化流水线中推断目标坐标。在六家主要 VRC 提供商（VTT、Geetest、NetEase、Dingxiang、Shumei、Xiaodun）上的评估中，VIPER 达到高达 93.2% 的成功率，在大多数基准上超越人类准确率。与已有求解器 GraphNet（83.2%）、Oedipus（65.8%）和 Holistic 方法（89.5%）相比，VIPER 始终优于所有基线。该框架在替代 LLM 后端（GPT、Grok、DeepSeek、Kimi）上同样保持鲁棒性，准确率维持在 90% 以上。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为了预判防御，我们进一步提出模板空间随机化（Template-Space Randomization，TSR），一种在不改变任务语义的前提下扰动语言模板的轻量级策略。TSR 可显著降低求解器（即攻击者）的性能。我们提出的设计为人类可解但机器难以破解的验证码指明了方向。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qi-minfeng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qi-minfeng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">37. Stayin&#39; Alive: How Global Stolen Data Markets Thrive on Telegram</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Tina Marjanov (剑桥大学); Taro Tsuchiya (卡内基梅隆大学); Konstantinos Ioannidis and Jack Hughes (剑桥大学); Nicolas Christin (卡内基梅隆大学); Alice Hutchings (剑桥大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">被盗数据是许多网络犯罪活动的催化剂，如垃圾邮件活动、鱼叉式网络钓鱼和身份盗窃。研究提供被盗数据的在线社区有助于打击这些犯罪活动。虽然匿名市场和论坛传统上是被盗数据的主要交易场所，但基于聊天的即时通讯应用 Telegram 已成为一种受欢迎的替代选择。鉴于 Telegram 对公众的可达性不断增强，被盗数据社区如何将其运营方式适配到该平台、规避审查努力并建立有韧性的社区，仍有待厘清。在本工作中，我们刻画了：i）被盗数据社区在 Telegram 生态系统中出现的位置；ii）它们提供的被盗数据类型；iii）它们的运营所在地；iv）它们如何逃避检测。本文提供四项主要贡献。首先，我们提供了迄今最大规模的 Telegram 被盗数据频道纵向数据集之一。在一年时间内，我们人工筛选了 1,521 个频道，收集了 1,400 万条消息和 360 万个共享文件。我们发现，被盗数据社区与 Telegram 上的其他社区在很大程度上是相互独立的。其次，我们对被盗数据的类型进行分类，旨在理解它们所助长的潜在网络犯罪。第三，已有文献聚焦于英语社区，而我们发现许多频道以非英语语言运营，并从非英语市场获取被盗数据。第四，这些社区部署了各种技术来规避监管。值得注意的是，提供指向其他被盗数据频道链接的&#34;网关频道&#34;在提升存活时长和增长率方面发挥着关键作用。最后，我们不仅为学术研究人员，也为不同司法管辖区寻求监控和审查此类活动的 Telegram 及执法机构提供了启示。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_marjanov.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_marjanov.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">38. Assumption-Free Fuzzy PSI via Predicate Encryption</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Erik-Oliver Blass (空客); Guevara Noubir (东北大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了首个高效的模糊隐私集合求交（Fuzzy Private Set Intersection，PSI）协议，它实现了线性通信复杂度，不依赖对参与方输入分布的严格假设，并且不使用低效的全同态加密。具体而言，该协议使两方能够计算各自集合中所有汉明距离在给定范围内的元素对，而对集合的结构不作约束。我们的关键洞察是，安全地计算两个输入之间的（阈值）汉明距离可以归约为安全地计算它们的内积。利用这一归约，我们使用近期关于内积谓词加密的技术构造了一个模糊 PSI 协议。为了在本文设定中使用谓词加密，我们证明了这些谓词加密方案仅需一种较弱的模拟安全概念。我们还展示了如何在无需可信第三方的情况下高效地分布其内部密钥派生。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">因此，建立在谓词加密之上的模糊 PSI 对任意输入分布实现了最优的线性通信复杂度。我们的实现验证了其可行性，并展现出相对于最密切相关工作的性能提升。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_blass.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_blass.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">39. Logos: Robust Sharding Blockchain With Fast Processing and Optimal Cross-Shard Overhead</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yizhong Liu (北京航空航天大学和北京未来区块链与隐私计算前沿创新中心); Boyu Zhao, Yuxuan Hu, Haojun Tan, Feiang Ran, Andi Liu, and Zhuocheng Pan (北京航空航天大学); Yuan Lu (中国科学院软件研究所); Song Bian, Jianwei Liu, and Zhenyu Guan (北京航空航天大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">分片区块链通过将网络划分为多个分片来显著提升可扩展性。由于涉及多个分片的跨分片交易（CSTX）占比可观，跨分片交易处理（CSTP）对系统安全与性能至关重要。然而，现有的 CSTP 方法在恶意节点通过无效 CSTX 泛洪引发鲁棒性受限方面存在不足，并带来较高开销，尤其在异步网络中更为突出。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出 Logos，一种具有快速 CSTP 和最优跨分片开销的鲁棒分片区块链。Logos 采用一种新颖的鲁棒广播—传输—共识模式。每个输入分片仅调用一个新设计的广播原语来生成输入可用性状态。在通过一种创新的并行单对单传输机制将这些状态交付给相关分片之后，有效 CSTX 通过共识协议被提交，而无效 CSTX 则被丢弃。我们证明 Logos 对有效 CSTP 实现了分片内最优开销，对无效 CSTP 实现了更低开销。此外，Logos 以最优跨分片开销实现可靠传输。在跨 4 个区域的 1000 个 AWS-EC2 节点上进行的实验表明，与基线（Kronos，NDSS&#39;25）相比，Logos 实现了 50% 的延迟降低，峰值吞吐量达 132.8 ktx/sec。此外，Logos 的跨分片网络使用量仅为 Kronos 的 1/210。在恶意泛洪攻击下，Logos 保持了 Kronos 2.86 倍的吞吐量，展现出强大的鲁棒性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-yizhong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-yizhong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">40. SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Mahbub Alam (德克萨斯农工大学); Muhammad Lutfor Rahman (加州州立大学圣马科斯分校); Sonjoy Kumar Paul, Amy W. Hays, Aftab Hussain, Md Imanul Huq, and Nitesh Saxena (德克萨斯农工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">钓鱼网站仍然是网络犯罪的主要推手。为此，许多学术界的基于 AI 的钓鱼网站检测方法被开发出来，它们常常启发了现实系统的设计。尽管大多数研究声称具有高准确率，但它们是否满足现实需求仍不清楚，例如对不断演变的钓鱼策略的抵御能力、在多样化良性页面上的鲁棒性、可解释性以及隐私性。我们提出 PHILTER（PHishing detection literature Inspection via LLMs and Targeted Expert Review，通过 LLM 和定向专家审查进行钓鱼检测文献审视），一个可扩展的框架，用于跨四项功能性指标和三项安全性指标对钓鱼网站检测研究进行定性评估。PHILTER 利用 LLM 提取证据并起草理由，随后由专家验证并据此产生最终评估。将其应用于 55 项学术方法后揭示了系统性缺陷。没有研究满足所有功能性和安全性需求。没有任何研究展现出有效应对多样化钓鱼策略的证据。大多数方法难以保护隐私并适应不断演变的攻击者策略，且许多方法因在多样化良性页面上测试有限而在实际中面临误报率升高的风险。我们还引入了检测策略的分类法（基于特征、基于相似度、基于身份和混合型），突出了设计权衡并有助于解释这些缺陷。我们的研究揭示了，以准确率为导向的评估会忽视削弱实际有效性的盲点，并揭示了一个关键开放挑战：在满足所有功能性和安全性需求的同时实现高准确率。我们提供了可操作的建议，以指导未来针对不断演变和自适应的钓鱼活动追求这一双重目标防御的设计。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_alam.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_alam.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">41. Streaming Function Secret Sharing and Its Applications</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Xiangfu Song (南洋理工大学); Jianli Bai (新加坡管理大学); Ye Dong (新加坡国立大学); Yijian Liu, Yu Zhang, and Xianhui Lu (中国科学院信息工程研究所和中国科学院大学网络空间安全学院); Tianwei Zhang (南洋理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">从软件和在线服务的用户处收集统计数据对于改善服务质量至关重要，但在保护个人隐私的同时获取这些洞察仍是一项挑战。函数秘密分享（Function Secret Sharing，FSS）是解决该问题的一种有前景的工具。然而，对于消息持续发送、安全计算任务反复在到达消息上执行的流式分析场景，基于 FSS 的解决方案仍面临若干挑战。我们引入一种称为流式函数秘密分享（streaming function secret sharing，SFSS）的新密码学原语，它是 FSS 的一种新变体，特别适合对流式消息的安全计算。我们对 SFSS 进行形式化，并提出具体构造，包括针对点函数、谓词函数的 SFSS，以及针对通用函数的可行性结果。SFSS 以简单且模块化的方式支撑若干有前景的应用，包括条件式转加密（conditional transciphering）、策略隐藏聚合和属性隐藏聚合。特别是，我们的 SFSS 形式化与构造识别出现有解决方案中的安全缺陷和效率瓶颈，而基于 SFSS 的解决方案以渐近和具体上都更优的效率和/或增强的功能实现了预期的安全目标。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-xiangfu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-xiangfu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">42. Shred-to-Shine Metamorphosis of (Distributed) Polynomial Commitments</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Weihan Li (北京航空航天大学网络空间安全学院; Ant Group); Zongyang Zhang (北京航空航天大学网络空间安全学院); Sherman S. M. Chow (香港中文大学); Yanpei Guo (新加坡国立大学); Boyuan Gao (北京航空航天大学网络空间安全学院); Xuyang Song (Anoma); Yi Deng (西安电子科技大学密码学院); Jianwei Liu (北京航空航天大学网络空间安全学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">简洁非交互式知识论证（SNARK）依赖多项式承诺方案（PCS）来简洁地验证多项式求值。基于线性码的高性能多线性 PCS（MLPCS）降低了证明者开销，而分布式 MLPCS 通过在证明者间并行化承诺与开启进一步削减了开销。采用快速 Reed–Solomon 交互式预言机近邻证明（FRI），我们提出 PIPFRI，一种将线性时间可编码码 PCS 的线性时间证明与 Reed–Solomon（RS）PCS 的紧凑证明和快速验证相结合的 MLPCS。通过降低快速傅里叶变换和哈希开销，PIPFRI 的证明速度比基于 RS 的 DeepFold（USENIX Security &#39;25）快 10 倍，同时保持有竞争力的证明大小和验证时间。与来自线性时间可编码码的 Orion（CRYPTO &#39;22）相比，PIPFRI 证明速度提升 3.5 倍，并将证明大小和验证时间降低 15 倍。作为一种线性可扩展的分布式变体，我们提出 DEPIPFRI，它增加了问责机制并将单个多项式分布在多个证明者之间，实现了首个面向通用电路的基于码的分布式 SNARK。值得注意的是，与缺乏问责机制且仅支持多个独立多项式的 DeVirgo（CCS &#39;22）相比，DEPIPFRI 将证明时间提升 25 倍，证明者间通信降低 7 倍。我们将 shred-to-shine 确定为关键洞察：将一个多项式划分为独立处理的片段，同时保持证明大小和验证时间。进入配对体制后，这一洞察产生了一种基于群的 MLPCS，其结构化参考串（SRS）缩短 16 倍，开启速度比 Kate–Zaverucha–Goldberg（TCC &#39;13）的多线性变体快 10 倍。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-weihan.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-weihan.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">43. Unbalanced Fuzzy Private Set Intersection for L_infinity Distance: Achieving Sublinear Communication with Large Set Size</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Shengzhe Meng and Xiaodong Wang (清华大学); Xv Zhou (北京航空航天大学); Bei Liang (北京数学科学研究院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">模糊隐私集合求交（Fuzzy PSI）是一种密码协议，使两方能够在近似匹配下计算各自集合的交集，其变体包括标准模糊 PSI 和带发送方隐私的模糊 PSI（PSI-SP）。尽管近期进展带来了高效的模糊 PSI 协议，但大多数针对双方集合大小相近的均衡情形设计。然而在实践中，许多应用涉及高度不均衡的集合（例如接收方集合远大于发送方集合，或反之）。本工作聚焦于 l∞ 度量下的不均衡模糊 PSI。我们观察到，现有协议的通信开销主要由遗忘型键值存储（OKVS）的传输主导，尤其是在集合大小不均衡时。如果 OKVS 是稀疏的，则可以使用批量隐私信息检索（BatchPIR）来降低这一开销。然而，此类优化需要具有特定属性的空间哈希，而现有空间哈希方案中鲜有满足这些要求者。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们重新表述空间哈希，并提出两种新方案：一种非交互式，一种交互式，各自适用于不同的输入集合条件。在此基础上，我们设计了两种不均衡模糊 PSI 协议，将稀疏 OKVS 与 BatchPIR 结合，在较大集合的大小上实现亚线性通信。第一个协议适用于接收方持有较大集合的场景，第二个则针对发送方拥有更多项的情形。我们的协议在通信和运行时间上显著优于现有最优方案。例如，在 100 Mbps 网络下，参数为 (N, M, d, σ) = (220, 25, 2, 10) 时，我们基于非交互式空间哈希的协议将通信从 16,128 MB（Baarsen and Pu，Eurocrypto&#39;24）降至 0.35 MB。此外，利用交互式空间哈希方法的模糊 PSI 协议在在线运行时间上至少快 31 倍，通信量降低 1762 倍，均相比 Gao 等人（Asiacrypt&#39;25）。对于带发送方隐私的模糊 PSI 协议，我们以至少 4 倍的在线运行时间提升和 6 倍的通信量降低优于 Piske 等人（CCS&#39;25）。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_meng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_meng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">44. mmCipher: Batching Post-Quantum Public Key Encryption Made Bandwidth-Optimal</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hongxiao Wang (香港大学); Ron Steinfeld (莫纳什大学); Markku-Juhani O. Saarinen (坦佩雷大学信息安全实验室); Muhammed F. Esgin (莫纳什大学); Siu-Ming Yiu (香港大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在安全群组通信和广播等应用中，高效地一次性向多个不同接收方交付多条消息非常重要。为此，多消息多接收方公钥加密（mmPKE）能够一次性为多个独立接收方批量加密多条消息，与逐条加密每条消息的平凡方案相比，显著降低了成本——尤其是带宽开销。这一能力在后量子设定中尤为理想，因为此时密文长度通常显著大于相应明文。然而，几乎所有先前关于 mmPKE 的工作都局限于量子易攻的传统假设。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们提出首个基于标准 Module Learning with Errors（MLWE）格假设的 CPA 安全 mmPKE 和多密钥封装机制（mmKEM），分别命名为 mmCipher-PKE 和 mmCipher-KEM。我们的设计分两步进行：（i）我们通过提出一种新的 PKE 变体——可扩展可复现 PKE（XR-PKE）——引入了一种新颖的 mmPKE 通用构造，该变体能够通过额外提示复现密文；（ii）我们使用一种新技术实例化基于格的 XR-PKE，该技术能够精确估计此类提示对密文安全性的影响，同时建立合适的参数。我们相信这两者都具有独立的研究价值。作为额外贡献，我们探索了抗自适应腐败和选择密文攻击的自适应安全 mmPKE 的通用构造。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们还提供了 mmCipher 的高效实现和实际性能的全面评估。结果表明，相对于现有最优方案，我们在带宽和计算上均有大幅节省。例如，对于 1024 个接收方，mmCipher-KEM 实现了 23–45 倍的带宽开销降低，密文仅比明文大 4–9%（接近最优带宽），同时计算成本降低 3–5 倍。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-hongxiao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-hongxiao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">45. DNS Cache Poisoning Like it&#39;s 2006</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Omer Ben-Simhon and Amit Klein (耶路撒冷希伯来大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">域名系统（DNS）支撑着几乎所有互联网服务，因此 DNS 解析的完整性对安全与可用性至关重要。我们对一类新型 DNS 缓存投毒攻击进行了全面研究，目标是最广泛部署的开源 DNS 解析器 BIND9。我们的攻击聚焦于两项关键能力，使其有别于大多数先前工作：（1）能够可靠地预测两个关键挑战参数——UDP 源端口和 TXID——而现有大多数攻击仅针对其中一个；（2）完全从客户端侧执行此预测，无需攻击者为其域名运行权威服务器，据我们所知这是首次。我们通过利用 BIND 伪随机数生成的弱点来实现这一点，即使在真实网络条件下也能实现高度可靠的预测。除纯客户端技术外，我们还开发了服务器端技术，以便攻击较旧的 BIND 9 的 9.18 分支。我们评估了这些攻击，并在多个 BIND 9 发布分支和配置中展示了实际的成功率。所有漏洞均已负责任地披露给 Internet Systems Consortium（ISC）和 FreeBSD Project，促成了两个补丁、CVE 编号及致谢。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ben-simhon.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ben-simhon.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">46. UncoreBleed: AEX-Free, High-Resolution, and Low-Noise Side-Channel Attacks on SGX Enclaved Execution</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Decheng Chen (华南理工大学); Zhi Zhang (西澳大利亚大学); Zhenkai Zhang (克莱姆森大学); Xin Zhang (山东大学); Yansong Gao (东南大学); Yi Zou (华南理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">可信执行环境如 Intel SGX 通过将 enclave 与操作系统和 hypervisor 隔离，提供强大的机密性和完整性保证。先前的工作声称 SGX 禁用 PMC 以缓解侧信道攻击。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文中，我们展示现代处理器具有 uncore PMC，其在 SGX 下的行为尚未被充分评估。利用这一观察，我们调查了生产模式 SGX enclave 中 PMC 的状态，并推翻了性能监控被抑制这一长期持有的信念：uncore PMC 记录与 enclave 执行相关的事件。我们进一步在 mesh-to-memory uncore 子系统中识别出一个关键事件，允许以 64 B 粒度进行基于地址的监控。通过逆向工程，我们揭示了其在支持 SGX 的 Xeon 处理器上的过滤机制、可编程性、可用性和地址映射。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于该事件，我们提出 UncoreBleed，首个基于 PMC 的、无 AEX、高分辨率且低噪声的针对 SGX 的侧信道攻击。UncoreBleed 能够从 enclave 内的 Libjpeg 重建图像，并在存在 TLBlur with AEX-Notify——现成 SGX 平台上最先进的软件防御——的情况下从单次解密中提取 RSA 私钥。我们的发现表明，处于活跃状态的 uncore PMC 对 enclave 机密性构成此前被低估的威胁，凸显了重新审视 SGX 关于性能监控安全假设的必要性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-decheng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-decheng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">47. M-Step: A Single-Stepping Framework for Side-Channel Analysis on TrustZone-M</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Cristiano Rodrigues (米尼奥大学 ALGORITMI 中心); Marton Bognar (鲁汶大学 DistriNet); Sandro Pinto (米尼奥大学 ALGORITMI 中心); Jo Van Bulck (鲁汶大学 DistriNet)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">可信执行环境（TEE）已成为将敏感 enclave 应用与不可信操作系统隔离的关键技术。针对 Intel SGX 和 TDX、AMD SEV 以及 Arm TrustZone-A 等高端平台的大量研究揭示了其在软件侧信道分析方面的局限，这些局限被利用特权定时器中断将 enclave 逐指令执行的专业单步进攻击框架所放大。与此同时，TEE 越来越多地部署于资源受限的物联网设备上，其中 Arm TrustZone-M 作为领先方案兴起，但其在高分辨率软件侧信道方面仍很大程度上未被探索。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出 M-Step，一个面向 TrustZone-M 的开放、可扩展的单步进攻击框架。尽管 Cortex-M 微控制器具有精确的定时器和确定性行为，但由于以下原因，实现精确的指令级步进仍具挑战：（i）缺少高端框架中使用的虚拟内存和页表；以及（ii）Cortex-M 独特的中断行为，某些多周期指令会被放弃或暂停以降低延迟。为克服这些挑战，我们广泛刻画了中断处理的 CPU 行为，并开发了一种新方法，利用此前被忽视的中断延迟泄漏来动态调整定时器中断。我们通过在最新的 Arm Mbed TLS 库中发现此前未知的漏洞来展示 M-Step 提升的分辨率和实用性，这些漏洞使单迹、确定性攻击能够从 TrustZone enclave 恢复完整 RSA 密钥。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_rodrigues.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_rodrigues.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">48. DaLens: Charting DNS Self-Amplification Threats at Large</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Liwen Xu and Zechao Cai (苏黎世联邦理工学院); Huayi Duan (香港科技大学(广州)); Adrian Perrig (苏黎世联邦理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">新兴的自我放大攻击（Self-Amplification Attacks，SAA）对域名系统（DNS）构成严重的拒绝服务（DoS）风险。它们能够大幅放大递归服务器与权威服务器之间的交互，以不成比例的低成本耗尽资源。评估此类攻击对全球名称解析基础设施的影响，对于 DNS 运营方有效分诊威胁并部署防御至关重要，然而这仍是一片未测绘的艰险领域。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们利用所设计开发的通用框架 DaLens，开展了首次大规模 SAA 测量研究。该工作包括梳理复杂的基础设施以识别有效的放大器，并以模块化、可扩展且可靠的方式量化其放大能力。在 307K 个持久公共解析器中，我们发现 29K 个唯一解析器集群可被并行用于 SAA，且其中相当数量即便相关漏洞已在先前工作中披露，仍能产生巨大的放大效应。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_xu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_xu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">49. Inconsistent, Incomplete, and Insecure: A Survey of Account Security Interfaces</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Arkaprabha Bhattacharya and Alaa Daffalla (康奈尔大学); Kevin Lee (独立研究员); Rosanna Bellini (纽约大学); Nicola Dell (康奈尔理工); Thomas Ristenpart (多伦多大学和康奈尔理工)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">尽管账户安全性有所改善，但账户被攻破的情况仍然普遍且危害严重，尤其是当攻击者在物理或社交上与受害者距离较近时（例如人际滥用场景）。为帮助用户识别未授权访问，网络服务提供账户安全接口（ASI）：通知和日志提供信息以帮助推断对抗性攻破。我们呈现迄今最大规模的 ASI 测量研究，评估了 100 个热门服务。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的研究凸显了一个令人不满的现状：29 个服务未向用户提供任何区分账户访问方式。在用新分类法对 ASI 进行分类后，我们表明各服务在所部署的类型上不一致。此外，ASI 通常不完整且令人困惑，即便对专家研究者也是如此。最后，在 61 个提供 ASI 以传达设备或位置描述的服务中，41 个（67.2%）易受欺骗攻击，使攻击者成功混淆访问来源。基于这些发现，我们提出六项改进未来 ASI 部署的原则。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bhattacharya.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bhattacharya.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">50. Hydrangea: Optimistic Two-Round Partial Synchrony with Improved Fault Resilience</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Nibesh Shrestha (Supra Research); Aniket Kate (Supra Research / 普渡大学); Kartik Nayak (杜克大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">部分同步设定下的共识协议面临一个基本权衡：实现最优拜占庭容错需要至少三轮的良好情况延迟，而在少于三轮内提交通常会降低容错能力。即便是乐观协议如 SBFT（DSN&#39;19）、FaB（TDSC&#39;06）和 Kudzu 在有利条件下实现了两轮的乐观良好情况延迟，但也仅以降低容错为代价。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们引入 Hydrangea，一种将低延迟与改进的容错相结合的部分同步拜占庭容错状态机复制协议。设 f 为所容许的最大拜占庭故障数，c 为所容许的最大崩溃故障数，k ≥ 0 为可调参数。对于一个由 n = 3f + 2c + k + 1 方组成的系统，当故障方总数（拜占庭或崩溃）至多为 p = c + ⌊k/2⌋ 时，Hydrangea 实现两轮的乐观良好情况延迟。在更具对抗性的设定下，即最多 f 个拜占庭故障和 c 个崩溃故障时，它保证三轮的良好情况延迟。我们进一步证明了一个匹配的下界：若 p &gt; c + ⌊(k+2)/2⌋，则在此故障模型下没有协议能在两轮内实现乐观提交。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们在地理分布式部署上的实验评估表明，无论在纯拜占庭故障模型还是拜占庭—崩溃故障模型下，Hydrangea 始终比现有最优协议实现显著更低的延迟，同时在吞吐量上也有适度提升。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shrestha.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shrestha.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">51. Revealing the Dark Side of Smart Accounts: An Empirical Study of EIP-7702 Incurred Risks in Blockchain Ecosystem</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Mingyuan Huang (香港科技大学); Han Liu (南开大学密码学与网络空间安全学院); Shuo Yang (中山大学); Daoyuan Wu (岭南大学); Shuai Wang (香港科技大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">EIP-7702 引入的智能账户代表了区块链账户抽象的重大进步，使外部拥有账户（EOA）能够在保留原始地址的同时升级为可编程账户。这一进步显著增强了账户功能和可用性，但也重新定义了 EOA 与智能合约账户（CA）之间的区块链信任边界，从而改变了安全假设并为新型攻击创造了机会。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为系统地审视这些风险，我们根据受害账户类型将基于智能账户的风险分为三类：针对 EOA、针对 CA 和复合攻击。随后，我们开发了结合大规模交易分析和跨合约静态分析的专用检测工具来识别恶意行为。将这些工具应用于支持 EIP-7702 的七条区块链，我们检测到 924 个恶意合约账户，其中包括若干此前未报告的零日案例。这些攻击已造成超过 230 万的损失，并使超过 1000 万暴露于潜在攻破风险。我们揭示了关于攻击者行为的多个关键洞察。具体而言，我们发现超过 63% 的 EIP-7702 授权交易与恶意 EOA 定向攻击相关，且最常被授权的合约中近一半由攻击者控制。此外，我们识别出攻击者用以规避检测的现有逃避策略、真实事件中观察到的攻击影响，以及未来部署中可能出现的潜在风险，凸显了在区块链生态中解决智能账户安全问题的紧迫性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-mingyuan.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-mingyuan.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">52. DMGuard: Safeguarding Kernels from Physical-Page Use-After-Free Vulnerabilities</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Juhee Kim, Jaeyoung Chung, Dae R. Jeong, and Byoungyoung Lee (首尔国立大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代内核依赖页表的完整性来实施高级安全措施。尽管这些防御已有效缓解了包括内存损坏在内的多种攻击，但攻击者已将目标转向破坏页表本身以绕过现有保护。异构地址转换域（包括独立的 CPU、GPU 和 IOMMU 页表）的兴起加剧了此类威胁，这些域对同步与一致性管理提出了更高要求。当虚拟地址仍映射到已被释放或重新分配的物理页时，攻击者可利用这一点访问任意物理内存。我们将此称为物理页释放后使用（physical-page use-after-free），区别于在虚拟地址上操作的传统堆释放后使用。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出 DMGuard，首个跨多种转换域全面应对物理页释放后使用漏洞的运行时缓解方案。DMGuard 采用轻量级无锁机制管理物理页状态机，确保页表中不存在悬空映射。在 Android 设备上的评估表明，DMGuard 以可忽略的性能开销有效阻断了所有已知的物理页释放后使用漏洞，证明了对新兴攻击的实用性与有效性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>暂未公开（论文处于禁运状态，PDF 将在 USENIX Security 2026 会议开幕首日发布）</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">53. Garuda and Pari: Faster and Smaller SNARKs via Equifficient Polynomial Commitments</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Michel Dellepere (独立研究者); Pratyush Mishra and Alireza Shirzad (宾夕法尼亚大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">SNARK 是强大的密码学原语，允许证明者为某个计算生成简洁证明。SNARK 研究的两个关键目标是最小化证明大小和最小化生成证明所需的时间。本工作提出了在两个目标上均突破前沿的新 SNARK 构造。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的第一个构造 Pari 是一种在所有已知 SNARK 中实现最小证明大小的 SNARK。具体而言，Pari 的证明大小仅为两个群元素和两个域元素，在以 BLS12-381 曲线实例化时总计仅 160 字节，小于 Groth16 [Groth, EUROCRYPT &#39;16] 和 Polymath [Lipmaa, CRYPTO &#39;24] 的证明大小。Pari 还实现了已知最低的链上 SNARK 验证 gas 成本，与 Groth16 相比降低 6%，与 FFLONK 相比降低 17%。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的第二个构造 Garuda 是一种通过首次支持任意“自定义”门和免费线性门（在密码学开销方面）来减少证明生成时间的 SNARK。这些优势使得与最先进的 SNARK 相比，证明者时间得到显著节省。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">两种构造都依赖于一种新的密码学原语：“等效率”多项式承诺（equifficient polynomial commitment, EPC）方案，该方案强制承诺多项式在特定基下具有相同的表示。我们为该原语提供了严格的安全定义，以及一元多项式和多重线性多项式的高效构造。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的构造通过一种新的编译器获得，该编译器将多项式 IOP 与我们的 EPC 方案相结合，得到简洁论证。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_dellepere.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_dellepere.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">54. Transparent Dictionaries from Polynomial Commitments</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hossein Hafezi (纽约大学); Alireza Shirzad (宾夕法尼亚大学); Benedikt Bünz and Joseph Bonneau (纽约大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出 IRONDICT，一种基于多项式承诺方案的透明字典构造。透明字典使不可信服务器能够维护可变字典，并可证明地向客户端提供查询服务。一个重大的开放性挑战是支持轻量级客户端的高效审计。已有方案要么服务器开销高（限制吞吐量），要么客户端查询验证开销大，阻碍了其在拥有数十亿用户的现代消息密钥透明性部署中的应用。我们的构造黑盒式地使用通用的多重线性多项式承诺方案，并继承其安全特性，即绑定性（binding）与零知识性（zero-knowledge）。我们使用近期的 KZH 方案实现该构造，发现在消费级笔记本上验证包含 10 亿条目的字典仅需 35 毫秒，比现有最优方案提升 300 倍。我们的构造还提供了小 150000 倍的证明（8 KB）和完美的隐私性，且客户端与服务器开销在具体实现上均高效。我们还展示了基于增量可验证计算（IVC）和检查点的快进技术，以实现更快的客户端审计。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hafezi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hafezi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">55. Efficient Threshold ML-DSA</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Sofía Celi (Brave Research); Rafael del Pino and Thomas Espitau (PQShield); Guilhem Niot (PQShield and 雷恩大学, 法国国家科学研究中心, IRISA); Thomas Prest (PQShield)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">门限签名方案允许一组用户联合生成数字签名，提供容错能力并增强去中心化。随着后量子密码学的兴起，基于格的门限签名作为可行替代方案受到关注。然而，现有构造在可扩展性、鲁棒性或与标准化方案的兼容性方面经常遇到挑战，特别是与 NIST 选定并标准化的基于模块格的数字签名算法（ML-DSA）的兼容性。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本工作提出了首个与 ML-DSA 完全兼容的门限签名方案，支持最多六方之间的安全高效签名。我们的构造利用先进的短秘密共享技术，并集成优化的拒绝采样，在分布式环境中实现了通信效率与正确性之间的良好平衡。我们用 Go 实现该构造，并在本地、局域网（LAN）和广域网（WAN）环境下评估其性能。基准测试表明，我们的门限 ML-DSA 方案不仅可实际部署，而且适用于现实应用场景，包括多设备加密货币钱包、基于门限的 TLS 认证以及 Tor 的目录权威机构。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_celi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_celi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">56. Nudge: A Private Recommendations Engine</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Alexandra Henzinger (麻省理工学院); Emma Dauterman (麻省理工学院 and 斯坦福大学); Henry Corrigan-Gibbs (麻省理工学院); Dan Boneh (斯坦福大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Nudge 是一种具有密码学隐私性的推荐系统。一个 Nudge 部署由三台基础设施服务器和大量用户组成，用户从大型数据集（如视频、帖子、商家）中检索/评分条目。Nudge 服务器周期性地以秘密共享形式收集用户评分，然后运行三方计算，在用户私有评分上训练轻量级推荐模型。最后，服务器向每个用户提供个性化推荐。在每一步中，Nudge 都不会向服务器泄露除聚合模型本身之外的任何用户偏好信息。用户隐私即使在攻击者攻破一台服务器全部秘密状态的情况下仍然成立。Nudge 的技术核心是一种新的三方矩阵分解协议。在拥有 50 万用户和 1 万个条目的 Netflix 数据集上，Nudge（在局域网中的三台 192 核服务器上运行）仅需 50 分钟即可私密地学习推荐模型，服务器间通信量为 40 GB。在标准质量基准（nDCG@20）上，Nudge 得分为 1.0 中的 0.29，与非隐私矩阵分解相当，仅略低于非隐私神经推荐器的 0.31。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_henzinger.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_henzinger.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">57. On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial Attack</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu, and Junhong Lai (浙江大学); Yaopeng Wang (东南大学 and 浙江大学); Kui Ren and Chun Chen (浙江大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大型视觉-语言模型（LVLMs）通过将视觉编码器的表示空间与 LLM 的表示空间对齐，在多模态任务中取得了显著成功。然而，它们仍易受可迁移对抗攻击的影响，这类攻击可在不访问模型的情况下操纵 LVLMs 的输出。因此，确保其可靠部署需要对黑盒鲁棒性进行严格评估。现有方法仅通过扰动 LVLMs 的视觉编码器进行有限的评估，且往往忽视无目标攻击场景。本工作提出模态对齐破坏攻击（Modality Alignment Breaking Attack, MABA），一种用于评估 LVLMs 黑盒鲁棒性的新型可迁移、无目标对抗攻击。MABA 强调破坏整个多模态流水线，针对两个关键阶段：视觉编码和模态对齐。首先，MABA 揭示可迁移对抗攻击的核心在于抑制判别性视觉表示，并明确将其作为优化目标以提高跨不同 LVLMs 的可迁移性。其次，MABA 引入一个互信息感知投影器，作为 LVLMs 模态对齐模块的替代，有效破坏跨模态一致性并增强可迁移性。大量评估表明，MABA 实现了最先进的性能，在图像描述任务中导致语义指标平均下降 58.37%。通过对多种 LVLM 家族的消融研究，我们获得了增强 LVLMs 鲁棒性的有价值洞察。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-zhichao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-zhichao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">58. From Easy to Hard++: Promoting Differentially Private Image Synthesis Through Spatial-Frequency Curriculum</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chen GONG and Kecen Li (弗吉尼亚大学); Zinan Lin (微软研究院); Tianhao Wang (弗吉尼亚大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">差分隐私（DP）合成图像通过在确保隐私保证的同时模仿敏感数据的统计特性，成为降低隐私担忧的关键工具。为提高合成图像质量，大多数研究聚焦于改进核心优化技术（如 DP-SGD）。近期出现了一种范式转变，即将这些技术作为现成工具，研究如何组合使用以获得最佳效果。其中一项 notable 工作是 DP-FETA，它提出使用“中心图像”来“预热”DP 训练，然后再使用传统 DP-SGD。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">受 DP-FETA 启发，我们探究是否还有其他工具可与 DP-SGD 组合使用。我们首先观察到，使用“中心图像”仅适用于包含大量相似样本的数据集。为处理图像差异较大的场景，我们提出 FETA-Pro，引入频率特征作为“训练捷径”。频率特征的复杂度介于空间特征（由“中心图像”捕获）和完整图像之间，可为 DP 训练提供更细粒度的课程。要将这两类捷径组合使用，一个挑战是处理空间特征与频率特征之间的训练差异。为此，我们利用生成模型的流水线生成特性（即无需一个模型同时训练多种特征/目标，而是使用多个模型分别处理不同特征，再将一个模型的生成结果输入另一个模型），采用更灵活的设计。具体而言，FETA-Pro 引入辅助生成器来产生与带噪频率特征对齐的图像，然后用这些图像连同空间特征和 DP-SGD 训练另一个模型。在五个敏感图像数据集上的评估表明，在隐私预算 ε = 1 下，FETA-Pro 比最佳基线平均保真度高 25.7%，效用高 4.1%。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">59. FIRA: Enabling Automatic Forensic Investigation of Unmanned Aerial Vehicles</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yizhi Huang (佐治亚理工学院); David Oygenblik (佐治亚理工学院); Runze Zhang, Mingxuan Yao, Muhammad Ibrahim, Burak Sahin, Haichuan Xu, Saman Zonouz, and Brendan Saltaformaggio (佐治亚理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在动态环境中，无人机（UAV）通常利用在线学习来优化其机器学习（ML）模型的决策边界，以提升性能。然而，当无人机变得不可恢复或不可用（如坠毁）时，取证调查人员将无法判断无人机的在线学习是否导致了坠毁。本文提出一种名为 FIRA 的新型取证技术，能够建立从 ML 模型到 UAV 系统组件的因果联系。FIRA 在飞行过程中回传在线学习更新和遥测数据（即使在带宽有限的情况下），并判断坠毁是否可归因于在线学习模型。我们将 FIRA 应用于 48 个 UAV 坠毁场景，使用两种广泛采用的 UAV 控制程序：PX4 和 ArduPilot。在四类 UAV 任务中，FIRA 各调查了 12 起由后门在线学习模型导致的坠毁事故，FIRA 能以 95.8% 的准确率正确地将模型归因为坠毁原因。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-yizhi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-yizhi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">60. MASLeak: Investigating and Exposing Intellectual Property Leakage Vulnerabilities in Multi-Agent Systems</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Liwen Wang (香港科技大学); Wenxuan Wang (中国人民大学); Shuai Wang, Zongjie Li, Zhenlan Ji, and Zongyi LYU (香港科技大学); Daoyuan Wu (岭南大学); Shing-Chi Cheung (香港科技大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大型语言模型（LLMs）的快速发展催生了多智能体系统（MAS），通过协作执行复杂任务。然而，MAS 的复杂性，包括其架构、智能体交互和复杂的内部通信处理，引发了关于知识产权（IP）保护的严重关切。本文介绍 MASLEAK，首个在实际黑盒环境下系统性地从 MAS 中提取 IP 的框架。我们假设一个现实的攻击者，只能向系统公共 API 提交查询并观察最终输出，对内部架构和后端 LLM 信息毫无先验知识。受计算机蠕虫传播并感染脆弱网络主机方式的启发，MASLEAK 精心构造对抗查询 q，以诱发、传播并保留来自每个 MAS 智能体的响应，从而揭示一整套专有组件，包括智能体数量、拓扑、系统提示、任务指令和工具使用。我们构建了首个包含 810 个 MAS 应用的合成数据集，并在真实 MAS 应用（包括 Coze 和 CrewAI）上评估 MASLEAK。MASLEAK 在提取 MAS IP 方面实现了高准确率，系统提示和任务指令的平均攻击成功率为 87%，在大多数情况下系统架构的攻击成功率为 92%。最后，我们讨论了这些发现的影响及潜在防御措施。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-liwen.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-liwen.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">61. Love, Lies, and Language Models: Investigating AI&#39;s Role in Romance-Baiting Scams</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Gilad Gressel and Rahul Pankajakshan (阿姆里塔大学 Amritapuri 校区网络安全系统与网络中心); Shir Rozenfeld (内盖夫本古里安大学); Ling Li (威尼斯大学); Ivan Franceschini (墨尔本大学); Krishnashree Achuthan (阿姆里塔大学 Amritapuri 校区网络安全系统与网络中心); Yisroel Mirsky (内盖夫本古里安大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">杀猪盘诈骗已成为全球范围内经济和情感伤害的重要来源。这些犯罪活动由有组织犯罪集团运营，将数千人贩卖为强迫劳动，要求他们在数周的文本对话中与受害者建立情感亲密关系，然后施压其进行欺诈性加密货币投资。由于此类诈骗本质上基于文本，它们引发了关于大型语言模型（LLMs）在当前和未来自动化中作用的紧迫问题。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们通过采访 145 名内部人员和 5 名诈骗受害者、开展一项将 LLM 诈骗智能体与人工操作者进行比较的盲法长期对话研究，以及对商业安全过滤器进行评估来调查这一交汇点。研究结果显示，LLMs 已在诈骗组织中广泛部署，87% 的诈骗劳动由易被自动化的系统化对话任务构成。在一项为期一周的研究中，LLM 智能体不仅从研究参与者那里获得了更大的信任（p=0.007），还比人工操作者实现了更高的请求遵从率（46% vs. 人工的 18%）。与此同时，主流安全过滤器对杀猪盘对话的检测率为 0.0%。综合来看，这些结果表明杀猪盘诈骗可能适合全面的 LLM 自动化，而现有防御仍不足以阻止其蔓延。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gressel.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gressel.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">62. United We Defend: Collaborative Membership Inference Defenses in Federated Learning</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Li Bai, Junxu Liu, Sen Zhang, Xinwei Zhang, Qingqing Ye, and Haibo Hu (香港理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">成员推理攻击（MIAs）通过判断特定数据点是否包含在目标模型的训练集中，对联邦学习（FL）构成了严重威胁。然而，现有的 MIA 防御通常独立应用于 FL 中的每个客户端，对于利用整个训练过程的时序信息来推断成员状态的强大基于轨迹的 MIA 效果不佳。本文研究一种由异构隐私需求和隐私-效用权衡驱动的新 FL 防御场景，其中仅防御部分客户端，以及客户端协作缓解成员隐私泄露的协作防御模式。为此，我们提出 CoFedMID，一种针对 FL 中 MIA 的协作防御框架，它限制本地模型对训练样本的记忆，并通过防御者联盟增强隐私保护和模型效用。具体而言，CoFedMID 由三个模块组成：用于选择性本地训练样本的类引导划分模块；回收有贡献样本并防止其过度自信的效用感知补偿模块；以及在联盟层面向客户端更新注入噪声以实现抵消的聚合中性扰动模块。在三个数据集上的大量实验表明，我们的防御框架显著降低了七种 MIA 的性能，同时仅带来很小的效用损失。这些结果在各种防御设置下均得到一致验证。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bai.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bai.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">63. Tracegram: Framing Trace-Level Traffic Analysis with Temporally-Aware Multiple Instance Learning</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jian Qu, Yuchen Zhang, Jialong Zhang, Jianfeng Li, and Xiaobo Ma (西安交通大学计算机科学与技术学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代网络行为跨越多个流并随时间演化，使得流间时序和共现上下文对于可靠的流量分析至关重要。这一需求在安全领域尤为关键，攻击在较长时间间隔和多个流中依次经历侦察、投递、命令与控制和横向移动阶段。现有的包级或单流方法割裂了这一上下文，限制了跟踪级分类、检测和归因的性能。我们将跟踪（trace）引入为分析单元，提出 Tracegram，将跟踪级分析形式化为多示例学习（Multiple Instance Learning）。Tracegram 将逐流编码器与时序感知聚合模块相结合，跨流推理，保留长程依赖，并产生支持分析人员验证和取证的流归因信号。我们的验证涵盖理论与实践。我们从理论上论证了基于 MIL 的跟踪级流量分析分解，并在四个公开数据集上跨多个任务进行了大量实验，表现出优于或可与最先进方法相媲美的性能。最后，对 DAPT 数据集中 APT 跟踪的案例研究表明，Tracegram 能突出与攻击阶段对齐的流，实现有针对性的调查。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">64. &#34;Abuse Risks are Often Inherent to Product Features&#34;: Exploring AI Vendors&#39; Bug Bounty and Responsible Disclosure Policies</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yangheran Piao, Jingjie Li, and Daniel W. Woods (爱丁堡大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">随着供应商采纳 AI 技术，安全研究人员正致力于发现和修复相关漏洞，鉴于 AI 系统处理敏感数据和关键功能，这一点至关重要。这一过程依赖于供应商接收并奖励 AI 漏洞报告。为评估当前实践，我们分析了 264 家 AI 供应商的漏洞披露政策。我们采用混合方法，结合快照式和纵向定性分析，并与 320 起 AI 事件和 260 篇学术文章进行对齐比较。分析显示，36% 的 AI 供应商没有建立相关政策，仅 18% 提及 AI 风险。数据访问、授权和模型提取漏洞最一致地被声明为范围内。越狱和幻觉最常被声明为范围外。我们识别出反映供应商对 AI 漏洞不同立场的三类画像：主动澄清型（n = 46）、沉默型（n = 115）和限制型（n = 103）。对齐结果表明，供应商在处理 AI 漏洞披露方面可能滞后于学术研究和真实事件。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_piao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_piao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">65. PICS: Private Intersection over Committed (and reusable) Sets</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Aarushi Goel (罗格斯大学); Peihan Miao and Phuoc Van Long Pham (布朗大学); Satvinder Singh (普渡大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">隐私集合交集（PSI）使两方能够在不泄露任何额外信息的情况下计算其私有集合的交集。虽然恶意安全 PSI 协议可防止多种攻击，但攻击者仍可通过在多次会话中使用不一致的输入来利用它们。这一局限源于安全多方计算中恶意安全的定义，但在 PSI 中尤为棘手，因为：（1）现实应用——如 Apple 用于 CSAM 检测的 PSI 协议和消息应用中的隐私联系人发现——通常需要对一致输入进行多次 PSI 执行；（2）PSI 功能使攻击者相对容易推断额外信息。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出基于承诺集合的隐私交集（Private Intersection over Committed Sets, PICS），一种通过承诺集在多次会话间强制输入一致性的新框架。在最先进的恶意安全 PSI 框架（即 VOLE-PSI [EUROCRYPT 2021]）基础上，我们使用轻量级密码学工具给出了 PICS 的高效实例化。我们的协议实现了强接收方输入一致性（即接收方使用确切的承诺集）和弱发送方输入一致性（即发送方无法向承诺集注入新元素，但可能使用承诺集的子集）。我们实现了该协议以展示具体效率。与 VOLE-PSI 相比，在集合大小 2^16-2^24 范围内，我们的通信开销为 1.57 - 2.04 倍的小常数，在各种网络设置下总端到端运行时间开销为 1.22 - 1.98 倍。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_goel.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_goel.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">66. MULCOTAINT: Towards Efficient Multi-tag Dynamic Taint Analysis via Hardware/Software Co-design</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Bing Qi (中国科学院大学; 中国科学院软件研究所); Yi Yang and Xiangkun Jia (中国科学院大学; 中国科学院软件研究所; 系统软件重点实验室（中国科学院）); Zhengpin Qian and Huafeng Huang (中国科学院大学; 中国科学院软件研究所); Purui Su (中国科学院大学; 中国科学院软件研究所; 系统软件重点实验室（中国科学院）)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">多标签动态污点分析（M-DTA）在漏洞分析等细粒度分析场景中至关重要。然而，当前软件方案存在严重的性能问题。硬件方案虽有前景，但仅支持单标签，难以扩展到 M-DTA。我们提出一种通过软硬件协同设计实现的高效 M-DTA 框架 MULCOTAINT。我们利用协处理器架构将污点分析与正常执行解耦，并解决了若干挑战，如将污点计算设计为向量化计算、用页表管理污点标签，以及提供污点分析引擎的功能接口。我们构建了包含 5 种类型 32 个程序的数据集，并进行了性能评估和漏洞分析实验。结果表明，MULCOTAINT 具有高性能和可接受的内存占用，且具备详细漏洞分析能力。MULCOTAINT 优于软件方案（TaintRabbit 和 PANDA）和硬件方案（HardTaint、RAFT 和 FineDIFT）。在各自基线上开销增长的最大差异为 MULCOTAINT vs. PANDA 的“1.14x vs. 4409.09x”，而 HardTaint 的平均开销增长是 MULCOTAINT 的 19.57 倍。尽管 MULCOTAINT 原型的硬件成本高于面向嵌入式的工作 RAFT 和 FineDIFT，但由于 M-DTA 逻辑复杂，该成本是可接受的。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qi-bing.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_qi-bing.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">67. WAVED: Principled Identification of Off-Path Exploitable Weak Verifications within the TCP/IP Protocol Suite</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yizhou Zhao and Xuewei Feng (清华大学); Min Li (中关村实验室); Ke Xu (清华大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">针对基础 TCP/IP 协议套件的离路（off-path）利用对互联网基础设施安全构成重大威胁。特别是，对接收载荷的弱验证——源于缺乏可靠信息进行验证或协议套件内的实现缺陷——导致的漏洞可被攻击者利用来操纵流量、引发数据丢失并中断受害服务器上的服务。本文首次系统研究这些漏洞，并提出 WAVED，一个用于识别 TCP/IP 协议套件实现中离路可利用弱验证的框架。WAVED 的核心是开发了针对 TCP/IP 内核的流敏感、上下文敏感和域敏感指针分析，并构建污点传播图（Taint Propagation Graph, TPG）来建模和追踪协议栈内的数据流。通过建模跨多种算术运算的字节粒度污点传播，我们的方法能精确定位与每个约束关联的特定输入字节。此外，计算方向敏感的污点信息以准确捕获和区分不同分支结果施加的约束强度，从而显著优于传统的字节不敏感和方向不敏感分析。我们在 IPv4 和 IPv6 上跨 Linux 5.15、Linux 6.8 和 FreeBSD 14.1 评估 WAVED。它精确揭示了 TCP/IP 中导致语义漏洞的弱验证，并发现了 14 个此前未知的漏洞。我们已向受影响的操作系统供应商负责任地披露了这些漏洞，并收到了 Linux 社区的致谢。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhao-yizhou.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhao-yizhou.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">68. CuSafe: Capturing Memory Corruption on NVIDIA GPUs</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hongyi Lu (南方科技大学 and 香港科技大学); Fengwei Zhang (南方科技大学); Zhenkai Zhang (克莱姆森大学); Shuai Wang (香港科技大学); Yanan Guo (罗切斯特大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代 GPU 应用，特别是在机器学习和科学计算领域，由于依赖 C/C++ 等内存不安全语言，越来越多地受到内存损坏漏洞的影响。然而，现有方案要么依赖商用 GPU 上不可用的硬件/软件，要么产生过高的性能开销，使其不适用于实际部署。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出 CuSafe，一种可直接部署在商用 NVIDIA GPU 上的新型 GPU 消毒器。CuSafe 采用将指针标记与带内缓冲区边界相结合的混合元数据方案，实现准确高效的内存安全验证。CuSafe 还引入了栈纪元追踪和虚拟地址随机化等机制，以缓解由时序损坏引起的元数据混淆。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们在 33 个程序上的安全评估表明，CuSafe 在现有 GPU 消毒器中独特地实现了空间漏洞和时间漏洞的最佳覆盖率。此外，我们在 44 个程序（包括 LLaMA2-7B 和 LLaMA3-8B 等大语言模型）上的性能基准测试显示，CuSafe 平均减速 13%，内存开销可忽略不计，仅 0.3%。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_lu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">69. Attacks on Approximate Caches in Text-to-Image Diffusion Models</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Desen Sun, Shuncheng Jie, and Sihang Liu (滑铁卢大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">扩散模型是一类强大的生成模型，能够根据用户提示生成图像及其他内容，但其计算开销巨大。为降低这一开销，近期的学术界与工业界工作采用了近似缓存（approximate caching）技术，即在一个缓存中复用来自相似提示的中间状态。该优化虽然高效，却因破坏了用户之间的隔离而引入了新的安全风险。本文对近似缓存所引入的安全漏洞进行了全面评估。首先，我们演示了利用近似缓存建立的一条远程隐蔽信道：发送方将带有特殊关键词的提示注入缓存系统，接收方即使在数天之后仍能恢复这些信息，从而实现信息交换。其次，我们提出了一种利用近似缓存的提示窃取攻击，攻击者可以根据缓存命中恢复已有的缓存提示。最后，我们提出了一种投毒攻击，将攻击者的标志嵌入先前窃取的提示中，导致命中被投毒缓存提示的请求出现意料之外的标志渲染。这些攻击均可通过服务系统远程执行，揭示了近似缓存中存在的严重安全漏洞。本工作的代码已公开。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_sun.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_sun.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">70. Side-Channel Attacks on Open vSwitch</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Daewoo Kim and Sihang Liu (滑铁卢大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">虚拟化技术在云系统中被广泛采用，以管理用户间的资源共享。虚拟化环境通常在宿主系统中部署虚拟交换机，使虚拟机之间以及虚拟机与物理网络之间能够通信。Open vSwitch（OVS）是最流行的软件虚拟交换机之一。它维护一个缓存层次结构，以加速从宿主机到虚拟机的数据包转发。我们从安全角度刻画了OVS内部的缓存系统，并识别出三种攻击原语。基于这些攻击原语，我们提出了三种通过OVS发起的远程攻击，破坏了虚拟化环境中的隔离性。首先，我们利用不同的缓存识别出远程隐蔽信道。其次，我们提出了一种新型的包头恢复攻击，能够泄露远程用户的包头字段，破坏了系统提供的机密性保证。第三，我们演示了一种远程数据包速率监控攻击，能够恢复远程受害者的数据包速率。为防御这些攻击，我们还讨论了潜在的缓解措施。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-daewoo.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-daewoo.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">71. A Distortion-minimization Watermarking Framework for Large Language Models: Larger Capacity, Stronger Robustness and Higher Quality</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Liming Zhai, Xuezhou Shang, Liyun Zhang, and Po Hu (华中师范大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）水印为生成文本提供可验证的来源识别，其实际部署需要大的水印容量、对攻击的强鲁棒性以及高质量的文本。然而，现有方法通常难以平衡所有这些指标，往往通过各自独立的设计来分别应对。为克服这一问题，我们提出了一种失真最小化水印（distortion-minimization watermarking, DMW）框架，在单一的优化范式中统一了容量、鲁棒性与质量。该框架将鲁棒性和质量建模为文本修改的失真代价，在给定水印长度下最小化总失真，从而实现最优权衡。具体而言，我们设计了多种失真代价：一种利用语义不变性来抵御攻击的鲁棒性代价，以及两种将修改引导至低内聚、高变异性区域以降低感知影响的质量代价。随后，我们提出了周期优化的校验子网格码（periodically optimized syndrome-trellis codes, PO-STCs），将整体失真最小化表述为一个周期性最短路径问题。这使得在序列生成过程中能够进行实时优化，并实现灵活的容量控制。在多个数据集和LLM上的大量实验表明，DMW在所有指标上均优于最先进的方法。值得注意的是，在严重的改写攻击下，DMW的匹配率比最佳基线高出46.35%，同时保持了更优的文本质量。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhai.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhai.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">72. Behind Bars: A Side-Channel Attack on NVIDIA MIG Cache Partitioning Using Memory Barriers</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Cheng Gu (罗切斯特大学); Reese Levine (加州大学圣克鲁兹分校); Zhenkai Zhang (克莱姆森大学); Tyler Sorensen (微软和加州大学圣克鲁兹分校); Yanan Guo (罗切斯特大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">NVIDIA多实例GPU（Multi-Instance GPU, MIG）是一项旨在为大型数据中心GPU提供隔离和安全多租户能力的功能。MIG将单个GPU划分为多个实例，每个实例拥有专用的硬件资源，如L2缓存分片。据文档记载，MIG还通过提供硬件隔离的可信执行环境，构成了NVIDIA机密计算栈的基础。然而，MIG的安全性声明值得更深入的调查，尤其是考虑到GPU内存系统的复杂性及其众多（文档稀疏的）内存指令。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们实证地研究了启用MIG时GPU L2缓存的行为。我们发现，尽管采用了分区设计，跨实例的L2缓存干扰仍然存在。具体而言，在一个MIG实例中生成的内存屏障（membar）会产生副作用，传播到其他L2分区，并影响其他实例中某些加载操作的时序。我们还发现，这些membar可由特定的GPU活动（如kernel启动）触发。基于这些观察，我们开发了一种新的基于时序的侧信道攻击，使一个MIG实例中的攻击者能够推断另一个实例中受害者的kernel启动模式。我们证明，该攻击危害了广泛使用的GPU应用（如大语言模型推理）的机密性，因为这些应用中的kernel启动模式与敏感信息相关。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gu-cheng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gu-cheng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">73. TrojPix: Electromagnetic Covert Channels via Imperceptible Pixel Modulation</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Guoming Zhang (山东大学和泉城实验室); Huiting Zhang, Zhenwei Lu, Heqiang Fu, Xin Gao, Riccardo Spolaor, and Yetong Cao (山东大学); Yanni Yang and Pengfei Hu (山东大学和泉城实验室)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">物理隔离网络依赖物理隔离来防止外部连接。此前的电磁（EM）隐蔽信道利用了视频线缆、内存总线和CPU的辐射，但很少能同时实现高吞吐量、长距离和视觉不可感知性，这限制了其在物理隔离环境中的实用性。我们证明，不可感知的像素调制可以确定性地在数字视频线缆上引发可控的电磁辐射，从而无需系统权限或硬件修改即可实现控制。基于这一发现，我们提出了TrojPix——一种在保持屏幕上不可感知性的同时，通过数字视频线缆实现高速、长距离通信的隐蔽信道。我们实现了一种轻量级通信方案，将像素到样本的映射与自适应解码相结合，在扩展距离上实现采样率级别的鲁棒通信。我们在真实条件下对九个商用现货（COTS）显示器制造商和十五条COTS数字视频线缆评估了TrojPix，在两种攻击模式（伪息屏和前景嵌入）下展示了其有效性。TrojPix实现了8.1 Mbps的峰值吞吐量和208 m的最大距离，揭示了对物理隔离网络安全的实用且隐蔽的威胁。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-guoming.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-guoming.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">74. Anonymous Tokens with Designated-Reader Metadata Bit</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Aisha Tu (武汉大学); Meng Jia (香港理工大学); Kun He, Jing Chen, and Ruiying Du (武汉大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">带有私有元数据位的匿名令牌在用户出示时向验证方传递隐藏信号，目前正处于标准化讨论中。现有方案仅允许令牌发行方读取信号，这给发行方带来了沉重负担，并且使得支持发行方隐藏变得困难，因为验证方必须联系发行方。在本文中，我们提出了一种带有指定读取者元数据位的匿名令牌方案，允许用户指定一个发行方接受的验证方直接从令牌中读取信号。我们还将方案扩展以支持读取者隐藏（对发行方和其他验证方隐藏用户所指定的验证方）以及发行方隐藏（防止验证方暴露令牌发行方）。我们证明了构造的安全性，并报告了其性能。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_tu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_tu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">75. Paper Title Under Embargo</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-style: italic;">（标题处于禁运状态，将在 USENIX Security 2026 会议开幕首日公开）</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Muyan Shen (中国科学院软件研究所; 中国科学院大学密码学院); Hongzhan Ma, Ketong Shang, and Ruofei Qu (中国科学院软件研究所); Yu Qin (中国科学院软件研究所, 北京, 中国); Dengguo Feng (中国科学院软件研究所)</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">76. TIMESLICE-SANDWICH: A GPU Side-Channel Attack Exploiting Time-Sliced Scheduling</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hodong Kim and Gyeongsup Lim (高丽大学); Seunghee Shin (纽约州立大学宾汉姆顿分校); Youngjoo Shin and Junbeom Hur (高丽大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代GPU支持并发应用之间的资源共享，由此引入了侧信道攻击的风险。虽然先前的研究已经探索了利用共享GPU资源的GPU侧信道，但时间片调度（当今GPU中资源共享的标准特性）在侧信道攻击方面的安全影响在很大程度上仍未被探索。在本研究中，我们分析了GPU时间片调度机制下并发执行引起的时序变化。我们首先确定时间片的上界，然后利用该上界来估计并发程序的时间片持续时间，最终使我们能够推断程序的整体GPU利用模式。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于这一发现，我们引入了TIMESLICE-SANDWICH，一种利用时间片持续时间变化来推断和区分受害者执行模式的新型GPU侧信道攻击。与先前的GPU侧信道攻击不同，TIMESLICE-SANDWICH不需要在特定共享资源上产生争用。在我们的实验中，TIMESLICE-SANDWICH在神经网络恢复攻击中平均达到94.40%的F1分数，在Google Chrome上的网站指纹攻击中平均达到92.84%的Top-1准确率，证明了其有效性。即使在存在噪声的情况下，我们的攻击在神经网络恢复中仍达到73.74%的平均F1分数。最后，我们讨论了针对现代GPU资源共享架构中由时间片模式引发的侧信道风险的潜在缓解措施。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-hodong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-hodong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">77. Bridging Usability and Performance: A Tensor Compiler for Autovectorizing Homomorphic Encryption</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Edward Chen, Fraser Brown, and Wenting Zheng (卡内基梅隆大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">同态加密（HE）通过支持在加密数据上进行计算，提供了强大的隐私保证。然而，HE中张量操作的性能高度敏感于明文数据打包到密文中的方式。大型张量程序引入了大量可能的布局分配，使得用户手动编写高效的HE程序既具有挑战性又枯燥乏味。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出了Rotom，一个将张量程序自动向量化为优化HE程序的编译框架。Rotom系统地探索了广泛的布局分配，应用了最先进的优化技术，并自动生成等效且高效的HE程序。其核心是，Rotom利用一种新型轻量级ApplyRoll布局转换算子，轻松修改底层数据布局，开启性能提升的新途径。我们的评估表明，Rotom能够在5分钟内可扩展地编译所有张量工作负载，将手工调优协议中的旋转次数减少最多3倍，并比先前的自动向量化系统实现最高80倍的性能提升。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-edward.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-edward.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">78. Shadowfax: Hybrid Security and Deniability for AKEMs</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Phillip Gajland (IBM苏黎世研究院); Vincent Hwang (马克斯·普朗克安全与隐私研究所, 拉德堡德大学); Jonas Janneck (波鸿鲁尔大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">随着密码学协议向后量子安全过渡，大多数协议采用了结合经典和后量子假设的混合方案。这一转变通常牺牲了效率、紧凑性甚至安全性。其中一个这样的性质是可否认性，它使用户能够合理地否认可能是定罪性消息的作者身份。虽然经典协议如X3DH密钥协商（用于Signal和WhatsApp）提供了可否认性，但后量子协议如PQXDH和Apple的iMessage（使用PQ3）则不然。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本工作通过研究如何在后量子协议中高效地保持可否认性来填补这一空白。具体而言，我们提出了两种用于认证密钥封装机制（AKEMs）的混合方案。第一种是黑盒构造，当两个组成AKEM均可否认时，能够保持可否认性。第二种是Shadowfax，一种非黑盒AKEM，实现了混合安全性，集成了经典非交互式密钥交换、后量子密钥封装机制和后量子环签名。Shadowfax在不诚实和诚实接收者设置中均满足可否认性，前者依赖统计安全性，后者依赖单个前量子或后量子假设。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">最后，我们提供了Shadowfax的多种可移植实现。当使用标准化组件（ML-KEM和Falcon）实例化时，Shadowfax的密文为1728字节，公钥为2036字节，在Apple M1 Pro上的封装和解封装开销分别为1.8M和0.7M周期。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gajland.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gajland.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">79. RBOOT: Accelerating Homomorphic Neural Network Inference by Fusing ReLU within Bootstrapping</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhaomin Yang, Chao Niu, Benqiang Wei, Zhicong Huang, Cheng Hong, and Tao Wei (蚂蚁集团)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">使用全同态加密（FHE）进行安全神经网络推理的一个主要瓶颈是ReLU等非线性激活函数的求值，这些函数在FHE下计算效率低下。最先进的解决方案使用高次多项式来近似ReLU，带来了显著的计算开销。我们提出了RBOOT，一个将ReLU求值无缝集成到CKKS自举（bootstrapping）中的优化框架，显著降低了乘法深度并提升了效率。我们的关键洞察是，CKKS自举中的EvalMod步骤由三角函数组成，而三角函数本身是非线性的。先前的工作将自举和激活函数视为独立的例程，错过了利用这种非线性的机会。通过协同优化这些组件，我们可以在自举过程本身中利用这种非线性来构造ReLU（及其他非线性函数），从而大幅减少计算开销。在四个广泛使用的CNN模型上的结果表明，与先前的多项式近似工作相比，RBOOT实现了2.77倍的端到端推理加速和81%的内存使用降低，同时保持了相当的准确率。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_yang-zhaomin.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_yang-zhaomin.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">80. PANGOLIN: Fuzzing Multilingual IoT Firmware with LLM-Driven Code Analysis</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhipeng Jia and Xiaokang Yin (信息工程大学); Shuitao Gan (先进计算与智能工程实验室); Chao Zhang (清华大学网络科学与网络空间研究院; JCSS, 清华大学（INSC） - 科学城（广州）数字科技集团有限公司); Hangtian Liu (数学工程与先进计算国家重点实验室); Jiangan Ji, Enzhou Song, and Ruijie Cai (信息工程大学); Jinglei Tan (数学工程与先进计算国家重点实验室); Shengli Liu (信息工程大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">多语言IoT通常指使用多种语言（如C、Python、Lua等）来实现其Web服务。虽然一些用户可访问的接口通过前端可视化以供交互，但在多语言IoT中，大量接口保持隐藏且未暴露给前端。此外，其参数通常呈现复杂的层级结构。从多语言设备中有效提取接口规范以进行漏洞发现是一个紧迫且尚未解决的问题。在本文中，我们提出了PANGOLIN，一种面向多语言IoT设备的新型模糊测试方案。首先，我们利用LLM分析API分发机制并识别接口。然后，我们引入一个LLM代理执行跨语言分析并生成输入参数规范。最后，我们利用响应驱动的反馈来修正参数规范。这些知识使语义感知的模糊测试能够探索更深的代码路径并发现更多漏洞。PANGOLIN成功发现了68个此前未知的漏洞，即比SOTA工具LABRADOR多2.96倍。值得注意的是，其中45个漏洞是在隐藏接口中发现的，而EAGLEYE仅能识别出4个此类案例。截至撰稿时，所有漏洞均已报告给厂商并被确认，已分配31个漏洞编号。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jia.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jia.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">81. Patch-Guided Vulnerability Detection: Extracting Java API Security Rules via Attack–Defense Cross-Analysis</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Bofei Chen, Shuang Liao, and Lei Zhang (复旦大学); Chibin Zhang and Mathias Payer (洛桑联邦理工学院); Yuan Zhang (复旦大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">安全敏感API是现代Java应用中的关键组件，然而对这些API的不当使用经常导致严重漏洞，如远程代码执行。现有的API安全规则生成方法存在局限性，因为它们依赖不完整的文档或基于发现的不一致性从源代码中推断模式。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了VulGenie，一个补丁驱动的框架，从已确认的安全补丁中提取精确的API安全规则，进而检测API误用漏洞。VulGenie解决了三个关键挑战。首先，它使用我们提出的新型修改行为依赖补丁图数据结构，从嘈杂的补丁中分离出被违反的约束和防御相关变更。其次，它识别受保护的安全敏感API，并通过攻防交叉验证合成规则。第三，它通过自适应的偏差引导静态分析扩展分析规模，以平衡精度和性能。在150个近期Java安全补丁上评估，VulGenie以81.82%的精度提取了198条API安全规则，发现了CodeQL中缺失的177条规则。在十个流行的Java应用上，VulGenie检测到46个0-day漏洞，大幅优于最先进的工作。通过我们负责任的漏洞披露，已有25个漏洞被修复，并分配了十个CVE编号。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-bofei.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-bofei.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">82. The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Louis Jannett (波鸿鲁尔大学); Andreas Mayer and Maximilian Westers (海尔布隆应用技术大学); Vladislav Mladenov (波鸿鲁尔大学); Christian Mainka (伍珀塔尔大学); Jörg Schwenk (波鸿鲁尔大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">通行密钥（Passkeys）提供了一种基于FIDO2和WebAuthn的安全且抗钓鱼的认证方法。它们近期获得了普及，越来越多的网站开始采用。然而，对这些网站进行大规模综合安全分析尚未得到充分解决。我们提出了PASSKEYS-RADAR，一个自2021年以来持续跟踪互联网上通行密钥部署的、不断更新的数据集。为构建该数据集，我们整合了多种来源，包括社区目录、Tranco 1M、CrUX 18M和历史互联网档案数据。我们分析了872个启用通行密钥的网站的收集数据，揭示了通行密钥的实现和管理方式。我们发现了网站在允许用户添加或删除通行密钥方面的重大差异，并发现网站要求认证器使用已弃用的加密算法。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为对启用通行密钥的网站进行全面安全评估，我们开发了PASSKEYS-ATTACKER。该工具允许在协议的每一步精确操纵WebAuthn消息，并集成了15种攻击类型，其中10种在先前工作中未被覆盖。其中，2种攻击类型具有严重的CVSS评分。我们在103个评估网站中的18个上发现了它们。这些攻击接管用户账户、删除其通行密钥或将其锁定在账户之外。近一半的测试网站（53个）至少容易受到一种高CVSS评分攻击的影响，使用户面临钓鱼和会话固定等威胁。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jannett.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jannett.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">83. StackWarp: Breaking AMD SEV-SNP Integrity via Deterministic Stack-Pointer Manipulation through the CPU&#39;s Stack Engine</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Ruiyi Zhang, Tristan Hornetz, Daniel Weber, Fabian Thomas, and Michael Schwarz (德国亥姆霍兹信息安全研究中心（CISPA）)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">机密虚拟机（CVM），如AMD SEV-SNP，旨在通过加密状态和约束特权控制来保护客户机操作系统免受不可信宿主机的侵害。这些平台承诺即使在同时多线程（SMT）保持启用的多租户云环境中也能提供隔离。虽然先前的攻击聚焦于内存层次结构或执行单元，但它们很大程度上忽略了前端配置。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出了StackWarp，一种利用AMD Zen CPU上的栈引擎在SEV-SNP客户机内修改栈指针的软件架构级攻击，完全破坏了完整性。StackWarp依赖于AMD Zen 1–5 CPU上一个共享的模型特定寄存器（MSR）中未公开的位，该位可启用或禁用栈引擎。我们的逆向工程表明，栈引擎的状态在逻辑核之间未正确同步，使攻击者能够跨Zen代际（包括已完全修补的Zen 5）在兄弟逻辑核上确定性地调整栈指针。我们通过对MSR空间（包括未公开的MSR）的系统性探索发现了StackWarp。通过翻转MSR位，我们发现了影响在兄弟逻辑核上运行的SEV-SNP客户机的位。为展示安全影响，我们在SEV-SNP客户机上展示了四种端到端攻击：RSA-CRT私钥恢复、OpenSSH密码认证绕过，以及使用sudo或内核态ROP链的权限提升。我们以软件加固指南作为总结，并主张在CVM活跃时进行微码或硬件更改以防止跨核控制栈引擎。我们的结果表明，当今保持SMT启用会破坏SEV-SNP的完整性保证。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-ruiyi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhang-ruiyi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">84. Leveraging Cryptographic Simulator Synthesis for Formally Verifying the FOO E-Voting Protocol</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>David Baelde (雷恩大学, 法国国家科学研究中心, IRISA); Adrien Koutsos and Justine Sauvage (法国国家信息与自动化研究所)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">密码学证明在很大程度上通过归约到以博弈表示的密码学假设来进行。这些归约依赖于模拟器，而模拟器的编写通常繁琐且包含大量平凡的代码。因此，在纸笔证明中模拟器仅被草拟，这容易出错。机械化密码学证明消除了出错风险，但要求用户显式编写模拟器是不合理的负担。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们考虑Squirrel中的模拟器合成问题，其中密码学模拟被表达为bi-deduction。虽然关于bi-deduction的开创性工作提供了一个证明系统和简单的证明搜索过程，但我们表明，在处理如IND-CCA2等博弈时它存在系统性失败。我们提供了一个显著改进的过程，能够在递归迭代中重用预言机调用，并生成精确的不变式来证明其正确性。我们在Squirrel中实现了该过程，并在FOO电子投票协议的选票隐私性证明中验证了它，这是FOO的首个计算性机械化证明，也是迄今为止最复杂的Squirrel证明。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baelde.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baelde.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">85. The Art of Hide and Seek: Making Pickle-Based Model Supply Chain Poisoning Stealthy Again</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Tong Liu and Guozhu Meng (中国科学院信息工程研究所; 中国科学院大学网络安全学院); Peng Zhou (上海大学); Zizhuang Deng (山东大学网络空间安全学院; 山东大学密码与数字经济安全国家重点实验室); Shuaiyin Yao and Kai Chen (中国科学院信息工程研究所; 中国科学院大学网络安全学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Pickle反序列化漏洞贯穿了Python的整个历史，虽广为人知但始终未获解决。由于其能够透明地保存和恢复复杂对象，许多AI/ML框架不顾其固有风险，继续采用pickle作为模型序列化协议。随着开源模型生态系统的增长，Hugging Face等模型共享平台吸引了大量参与者，显著放大了pickle利用的现实影响，并为模型供应链投毒开辟了新途径。虽然已有多种最先进的扫描器被开发用于检测投毒模型，但它们对投毒面不完整的理解使攻击者能够绕过它们。在本工作中，我们首次从模型加载和危险函数两个角度系统性地披露了基于pickle的模型投毒面。我们的研究展示了基于pickle的模型投毒如何保持隐蔽性，并揭示了当前扫描方案的关键缺口。在模型加载面，我们在五个基础AI/ML框架中识别出22条不同的基于pickle的模型加载路径，其中19条被现有扫描器完全遗漏。我们进一步开发了一种名为异常导向编程（Exception-Directed Programming, EDP）的绕过技术，发现了9个EDP实例，其中7个能够绕过所有扫描器。在危险函数面，我们发现了133个可利用的gadget，实现了几乎100%的绕过率。即使面对表现最佳的扫描器，这些gadget仍保持89%的绕过率。通过系统性揭示基于pickle的模型投毒面，我们实现了对现实扫描器的实用且鲁棒的绕过。我们向相应厂商负责任地披露了发现，获得了确认和12,000美元的漏洞赏金。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-tong_0.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-tong_0.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">86. B-Privacy: Defining and Enforcing Privacy in Weighted Voting</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Samuel Breckenridge, Dani Vilardell, and Andrés Fábrega (康奈尔科技学院, IC3); Amy Zhao (Ava Labs, IC3); Patrick McCorry (Arbitrum基金会); Ari Juels (康奈尔科技学院, IC3)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在传统的“一人一票”投票系统中，隐私等同于选票保密性：投票统计结果会被公布，但单个投票者的选择被隐藏。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">然而，按持币比例对投票进行加权的投票系统如今在加密货币和 web3 系统中已十分普遍。我们表明，这些加权投票系统颠覆了既有的投票者隐私概念。我们的实验表明，即使在选票保密的情况下，公布原始统计结果也往往会暴露投票者的选择。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">因此，加权投票需要一种新的隐私框架。我们引入了一个称为 B-privacy 的概念，其基础是贿赂——当今投票系统中的一个关键问题。B-privacy 刻画了对手基于已公布的投票统计结果贿赂投票者所需付出的经济成本。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了一种通过对投票统计结果加噪来提升 B-privacy 的机制。我们证明了其在 B-privacy 与透明度（即所报告统计结果的准确性）之间权衡的界。我们在 27 个去中心化自治组织（DAO）的 2,503 项提案上进行了实验，结果表明，在透明度几乎不下降的情况下，我们的机制将 B-privacy 提升了 3.5 倍的几何平均因子。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的工作为加权投票系统提供了首个有原则的、实用的、系统性的指导，补充了现有以选票保密性为核心的方法。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_breckenridge.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_breckenridge.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">87. InstrSem: Automatically and Generically Inferring Semantics of (Undocumented) CPU Instructions</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Lorenz Hetterich, Fabian Thomas, Tristan Hornetz, and Michael Schwarz (德国亥姆霍兹信息安全研究中心（CISPA）)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代 CPU 实现了复杂的指令集架构（ISA），但机器可读的语义往往并不完整。更糟糕的是，许多 CPU 支持未文档化的指令，即那些能在硬件上执行但未出现在规范中的比特串，这可能导致潜在的安全漏洞。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出 InstrSem，一种与 ISA 无关的、模块化的、全自动化的方法，仅从执行行为推断指令语义，并提供人类和机器均可理解的语义。从原始编码出发，InstrSem 在系统变化的各种架构状态下执行该编码，并综合出能够解释每个状态分量变化的紧凑数学函数。通过对编码比特进行变异，并将引发的行为变化与比特位置相关联，InstrSem 随后从单个编码泛化到完整指令，恢复寄存器和立即数字段。与以往专注于单一 ISA 的工作不同，InstrSem 是通用的。它仅需一个轻量级 ISA 模型和每种架构一个用户态运行器，并支持定长和变长编码（RISC 和 CISC）、内存访问以及条件行为。我们在 RV64I、AArch64 和 LA64 上评估了 InstrSem，并额外在 Logitech 宏语言和部分 x86-64 上展示了 CISC 适用性。InstrSem 自动恢复了 RV64I 基础指令集中超过 97.81% 的正确语义，并在 77 小时内恢复了 LA64 指令集中覆盖 1,009,055,744 种指令编码的 136 条指令的语义。InstrSem 发现了未文档化的向量指令、QEMU 与 Loongson 硬件之间的不一致性，以及会使 QEMU 崩溃的指令。InstrSem 实现了指令语义的可扩展恢复，大幅自动化了商品级和冷门目标的逆向工程，并为仿真、验证和安全分析奠定了更坚实的基础。凭借支持新架构的极低要求、模块化设计以及人类可读的输出，InstrSem 能够辅助未来的安全分析。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hetterich-lorenz-instrsem.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hetterich-lorenz-instrsem.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">88. Paper Title Under Embargo</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-style: italic;">（标题处于禁运状态，将在 USENIX Security 2026 会议开幕首日公开）</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Lorenz Hetterich, Tristan Hornetz, Fabian Thomas, and Michael Schwarz (德国亥姆霍兹信息安全研究中心（CISPA）)</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">89. Sliding into the Flight Deck&#39;s DMs: Practical Message Attacks on CPDLC</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Mehdi Ziazi (苏黎世联邦理工学院); Khalid Aleem (独立研究者); Harshad Sathaye (苏黎世联邦理工学院); Martin Strohmeier (网络防御园区，armasuisse科学与技术)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Controller–Pilot Data Link Communications（CPDLC）系统已成为现代空中交通管理不可或缺的组成部分，尤其是在语音通信受限或不可用的高密度或海洋空域中。CPDLC 旨在提高运行效率，是传统甚高频（VHF）语音通信的替代方案，采用标准化的数字消息来传达高度变更、航向调整、自由文本消息和频率切换。然而，CPDLC 并未实现加密，主要依赖协议的复杂性和模糊性来防止被滥用。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本研究中，我们对 CPDLC 进行了全栈安全分析，展示了若干漏洞，这些漏洞允许通过伪造地面站攻击劫持 ATC-飞行员链路，并实施大规模拒绝服务攻击，可使无线电范围内的所有飞机的 CPDLC 服务瘫痪。作为概念验证，我们还推出了 cpdlc-gs，这是首个基于 SDR 的全栈 CPDLC 地面站实现，能够注入上行链路消息以发送虚假 CPDLC 飞行指令并实施有效的拒绝服务攻击。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">此外，为了评估 cpdlc-gs，我们与空中导航服务提供商和航空电子设备制造商合作，利用 Universal Avionics 的真实可认证硬件，开发了一个全新的、功能完备的测试环境。通过这一设置，我们构思并验证了若干攻击，证明即使是孤立的伪造地面站也可能构成重大威胁，尤其是当飞行员处于高工作负荷或通信降级场景时。总体而言，我们认为 CPDLC 的广泛依赖与全球普及使其成为高价值目标，而滞后的航空数据链安全标准化进程亟待解决。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>暂未公开（论文处于禁运状态，PDF 将在 USENIX Security 2026 会议开幕首日发布）</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">90. You Know Why, but Still Rely: The Impact of Explainable AI on Trust, Task Load, and Performance in Cybersecurity Decision-Making</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Neele Roch, Hannah Sievers, Noé Zufferey, and Verena Zimmermann (苏黎世联邦理工学院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">随着机构数字化的不断推进，对有效网络安全措施的需求迅速增长。与此同时，网络安全任务的复杂性和数量正超越现有从业者的处理能力。利用 AI 增强人类网络安全专长有望降低复杂性和认知过载。对 AI 决策的透明且人类可理解的洞察，不仅是欧盟等治理机构的要求，也是从业者在高风险场景中与 AI 协作时自身的需求。我们报告了一项被试间研究（N = 139），考察了可解释 AI（XAI）的解释对具有网络安全领域知识的用户在恶意域名拦截场景下的信任、可用性、感知任务负荷和协作任务绩效的影响。在该场景中提供解释并未促进信任；事实上，具有领域知识的用户在与 XAI 交互后报告了更低的信任。定性结果表明，他们会应用自身的决策标准，而暴露 AI 的决策边界可能引入歧义并助长不信任。尽管纳入 XAI 未增加感知任务负荷，但也未能提升绩效。这些发现对当前 XAI 方法在以知识为中心的决策场景中的有效性提出了重要问题，并凸显了在网络安全领域需要更具情境敏感性、与用户对齐的解释策略。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_roch.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_roch.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">91. Silicon Heist: (Ransom) Attacks for Cloud FPGAs via Privilege Escalation</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Simon Klix, Felix Hahn, Maik Ender, Nils Albartus, and Christof Paar (马克斯·普朗克安全与隐私研究所（MPI-SP）); Russell Tessier (马萨诸塞大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于云的 FPGA 已成为价值数十亿美元的产业，允许用户借助云基础设施的可扩展性和灵活性部署自定义硬件设计。在云服务提供商（CSP）拥有的硬件上运行用户设计会引入风险，包括蓄意的硬件损坏和针对主机的拒绝服务（DoS）攻击。为缓解这些风险，CSP 实施了限制用户设计并防止未授权行为的安全机制。我们提出了一条在 AMD FPGA 上的新型提权路径，利用（i）内部配置访问端口（ICAP）绕过提供商防御，（ii）逐步将攻击者能力提升至远程 JTAG 访问，并（iii）调查由此产生的威胁向量。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">任何普通云客户都可以恶意获取此类 ICAP 访问权限，从而无限制地重新配置部分 FPGA 逻辑资源——使传统的云 FPGA 攻击得以复现。通过 ICAP，用户最终可以获得对硬件低层 JTAG 接口的远程控制，从而访问器件的 eFuse。这种访问反过来允许攻击者不可逆地编程加密设置，从而禁用未来的重新配置，并将 CSP 锁在其自有设备之外。攻击者可以利用这种提升的权限实施勒索软件攻击，云提供商必须支付赎金以换取解密密钥才能重新控制其设备——这实际上引入了首个针对 FPGA 的勒索软件。在研究了这一新型提权路径之后，我们在亚马逊的 EC2 F1 和 F2 实例上验证了其可行性，并探讨了所启用攻击向量的影响。我们由此揭示了云环境中未受保护的低层硬件组件被忽视的威胁。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_klix.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_klix.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">92. Bridging Bitcoin to Second Layers via BitVM2</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Robin Linus Woll (斯坦福大学和 ZeroSync 协会); Lukas Aumayr (爱丁堡大学和 Common Prefix); Zeta Avarikioti (维也纳工业大学和 Common Prefix); Matteo Maffei (维也纳工业大学); Andrea Pelosi (比萨大学、卡梅里诺大学和维也纳工业大学); Orfeas Stefanos Thyfronitis Litos (伦敦帝国理工学院和 Common Prefix); Christos Stefo (维也纳工业大学); David Tse (斯坦福大学和 Byzantine Research); Alexei Zamyatin (BOB)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">区块链基础设施中的一个圣杯是比特币与其二层网络或其他链之间的无须信任的桥。我们通过引入首个基于轻客户端的比特币桥在这一愿景上取得进展。其核心是 BitVM2-CORE，一种新型范式，能够在比特币上实现任意程序执行，将图灵完备的表达能力与比特币共识的安全性相结合。BitVM2-BRIDGE 推进了以往的方法，在设置阶段将信任假设从诚实多数（t-of-n）降低到存在性诚实（1-of-n）。仅需一个理性操作者即可保证活性，且任何用户都可以充当挑战者，实现无许可验证。我们已经开发了 BitVM2 的生产级实现，并在比特币主网上执行了完整的挑战验证。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_woll.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_woll.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">93. &#34;Oh, what people would do with my knife?&#34; Navigating the Dual-Use Dilemma in PoC Exploit Development, Disclosure, and Community Dynamics</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Arwa Al Alsadi and Lorenz Kustosch (代尔夫特理工大学); Lamya Alowain (独立研究者); Michel Van Eeten and Carlos H. Gañán (代尔夫特理工大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">随着概念验证（PoC）漏洞利用在披露后数分钟内就从演示转变为武器化攻击，网络安全领域面临着日益严峻的挑战。尽管已有研究记录了其时间动态和恶意部署，但在理解 PoC 创建背后的人为因素方面仍存在关键空白。通过对不同地区的 16 位 PoC 开发者进行半结构化访谈，我们应用期望-价值理论揭示了 PoC 开发是一个复杂的动机生态系统，技术信心、价值评估和风险计算在双重用途的张力中交织。我们证明 PoC 开发涵盖从崩溃演示到武器化漏洞利用的连续谱，由多方面的权衡而非二元伦理所塑造。我们识别出三个理论延伸：使责任外部化的双重用途道德推理、厂商行为重塑披露决策的动态价值评估，以及在伦理研究与技术精通之间的身份导航。厂商响应性、社区动态和法律约束显著影响披露策略。PoC 开发者在应对安全改进与潜在滥用之间的张力时采用风险缓解方法，这挑战了对“负责任”与“不负责任”披露的二元划分。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_al-alsadi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_al-alsadi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">94. Lost in Blockchain Address Misuse: Hidden Cross-Platform Risks and Their Security Impact</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhenzhe Shao (中山大学和浙江大学); Jiashuo Zhang (北京大学); Zihao Li (电子科技大学和香港理工大学); Daoyuan Wu (岭南大学); Chong Chen and Yiming Shen (中山大学); Lingfeng Bao (浙江大学和杭州高新区（滨江）区块链与数据安全研究院); Yanlin Wang (中山大学); Jiachi Chen (浙江大学和杭州高新区（滨江）区块链与数据安全研究院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">以太坊等区块链系统采用基于账户的模型，其中每个账户由地址唯一标识。地址作为用户交互和资产安全的基础接口至关重要，但在被误用时也会带来重大风险。在本文中，我们系统地揭示并分析了一类称为“地址误用”的风险，包括两大类别：合约账户（CA）误用和外部拥有账户（EOA）误用。具体而言，当用户错误地将非合约地址（NCA）当作 CA 对待时即发生 CA 误用，而当用户与私钥已暴露的 EOA 交互时则发生 EOA 误用。对于每一类别，我们揭示了其底层机制，并引入了此前未公开的攻击向量，使攻击者能够利用这些漏洞获利。为评估其普遍性和影响，我们首先从 GitHub 和 Stack Exchange 构建了一个数据集，其中包含各种区块链网络的地址。该数据集包含 1000 万个用于误用分析的候选地址和 1600 万个已暴露的私钥。然后，我们在以太坊和 BSC 上对其关联交易进行大规模链上分析。通过结合启发式规则、交易模式分析和符号执行，我们识别出 65,340 个高风险地址实例，关联的资产损失约达 127k ETH 和 17.7k BNB，相当于超过 5.748 亿美元。我们评估了检测方法的准确性以确保结果的可靠性，总体精确率达到 99.11%。此外，我们的实证评估还揭示了两个此前未公开的新攻击向量，提供了攻击者如何积极利用用户地址误用获利的真实证据。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">95. DDR-SSE: Duplicated Retrieval of Documents for System-wide Secure Searchable Symmetric Encryption</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zichen Gui (美国佐治亚大学); Simon-Philipp Merz and Kenneth G. Paterson (瑞士苏黎世联邦理工学院); Sikhar Patranabis (IBM印度研究院)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">可搜索对称加密（SSE）方案能够在加密文档上进行高效的关键词搜索，代价是泄露部分信息。如果一个 SSE 方案能够抵御可访问加密索引和加密文档检索泄露的对手的密码分析，则称其为系统级安全的。绝大多数最先进的 SSE 方案实际上都不是系统级安全的（Gui 等，IEEE S&amp;P 2023）。目前，唯一高效且系统级安全的 SSE 方案是 SWiSSSE（Gui 等，PoPETS 2024）。然而，SWiSSSE 要求客户端状态在每次查询时更新（这阻碍了在各种实际场景中的采用），且其泄露难以精确刻画（从而使安全分析更困难）。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出 DDR-SSE——一种实用高效、系统级安全的 SSE 方案，仅需静态客户端状态，且具有简单的泄露特征。在技术上，我们引入了一种新型加密文档检索方案，利用重复文档存储和随机化文档检索来抑制访问模式泄露，同时不牺牲实际效率。该方案的一个显著特点是其概念上的简洁性（不同于 SWiSSSE 使用了极其复杂的文档检索机制）。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们针对严格形式化的系统级泄露特征给出了 DDR-SSE 的基于模拟的安全证明。通过广泛的泄露密码分析，我们证实 DDR-SSE 对查询重构攻击具有韧性（即使在“不切实际地”强的攻击假设下）。最后，我们对 DDR-SSE 的原型实现进行了基准测试，表明它能平滑扩展到真实应用中所见规模的大型数据库。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gui.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gui.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">96. Concretely Efficient Blind Signatures Based on VOLE-in-the-Head Proofs and the MAYO Trapdoor</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Carsten Baum and Marvin Beckmann (丹麦技术大学); Ward Beullens (IBM，苏黎世); Shibam Mukherjee (格拉茨工业大学和格拉茨知识中心); Christian Rechberger (格拉茨工业大学和 TACEO，格拉茨)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">盲签名（Chaum，CRYPTO 82）是许多隐私保护应用（如匿名凭证或电子现金方案）中的重要构建模块。近年来，基于后量子假设（主要是格）构建盲签名引起了浓厚兴趣。虽然性能已有改善，但在计算和通信方面尚无构造达到实用效率。当前最先进的方法在每次向验证者展示基于格的盲签名时，至少需要 20 KB 的通信量，且证明者时间超过 100 ms。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们提出了一条替代方向，即一个具有后量子合理性的盲签名方案 PoMFRIT。它构建于 VOLE-in-the-head 零知识证明系统（Baum 等，CRYPTO 2023）之上，我们将其与 MAYO 数字签名方案（Beullens，SAC 2021）相结合。我们实现了 PoMFRIT 的多个版本以展示安全性与性能的权衡，并提供了构造的详细基准测试。签名签发对大小为（6.7）KB 的盲签名需要（0.45）KB 的通信量。即使对于具有 128 位安全性的保守构造，展示盲签名也可在 76 ms 以内完成。作为盲签名方案的构建模块，我们实现了首个用于 SHA-3 系列哈希函数的 VOLE-in-the-head 证明，我们认为这具有独立的研究价值。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baum.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baum.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">97. Turn Your Face Into An Attack Surface: Screen Attack Using Facial Reflections in Video Conferencing</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yong Huang, Yanzhao Lu, Mingyang Chen, En Zhang, and Jiazi Li (郑州大学); Wanqing Tu (杜伦大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在视频会议中，人脸是主要的视觉焦点，发挥着增强视觉交流和情感连接的多种作用。然而，我们认为人脸也是一个侧信道，可能在在线视频流中不知不觉地泄露屏幕上的信息。为此，我们进行了可行性研究，结果表明，在环境光和显示器发出的光线照射下，人脸能够反映不同屏幕内容的光学变化。随后，本文提出 FaceTell，一种新型侧信道攻击系统，可在视频会议期间从普遍却微妙的面部反射中窃听细粒度的应用活动。我们在一个真实测试平台上实现了 FaceTell，使用了三个不同品牌的笔记本电脑和四个主流视频会议平台。随后用 24 名受试者在 13 个独特的室内环境中对 FaceTell 进行了评估。凭借超过 12 小时的视频数据，FaceTell 在窃听 28 个流行应用时达到了 99.32% 的高准确率，并对许多实际影响因素具有鲁棒性。最后，我们提出了潜在的对策以缓解这种新型攻击。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-yong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-yong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">98. PROBE+DETECT+MITIGATE (PDM): Enabling Cloud Tenants to Self-Defend against Microarchitectural Attacks</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Arash Daneshmand (不列颠哥伦比亚大学奥卡纳根分校和康考迪亚大学); Hugo Kermabon-Bobinnec (康考迪亚大学); Lingyu Wang (不列颠哥伦比亚大学奥卡纳根分校和康考迪亚大学); Makan Pourzandi (爱立信安全研究院，爱立信加拿大); Suryadipta Majumdar (康考迪亚大学); Yosr Jarraya (爱立信安全研究院，爱立信加拿大)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">微架构攻击是公有云环境中的一个关键安全问题，因为它们可能导致利益冲突的云租户之间发生信息泄露。现有解决方案通常需要提供商级资源，如硬件性能计数器或主机进程，这些可能对云租户不可用。云租户缺乏意识可能促使云提供商推迟部署厂商补丁，正如 PRIME+PROBE 和 Spectre 变种等已打补丁但仍活跃的威胁所证明的那样。在本文中，我们提出 PDM，一种使云租户能够独立检测和缓解微架构攻击而无需提供商帮助的解决方案。首先，PDM 引入了基于租户的检测，其基于一个有趣的观察，即使用流行的 FLUSH+RELOAD 攻击技术探测受害者应用的内存空间实际上可以用于检测。其次，PDM 通过在检测时选择性地触发混淆和内存内加密技术来实现高效的基于租户的缓解。第三，我们解决了若干关键挑战，包括（i）不涉及驱逐的攻击（如 Spectre），（ii）对源代码或二进制插桩的需求，（iii）来自受害者或同驻租户的良性噪声，以及（iv）准确性、延迟和开销之间的权衡。我们的实验表明，PDM 使租户能够准确（例如，在我们的测试平台上 TPR ≥99.72% 且 FPR ≤0.13%，在 AWS Fargate 上 TPR ≥98.63% 且 FPR ≤0.83%）、及时（例如，触发缓解有 7ms 的提前时间）、高效（例如，在 SPEC CPU 2017 上开销 ≤2.47%）且鲁棒（对噪声和规避性攻击均如此）地检测和缓解各种微架构攻击，包括 PRIME+PROBE 和 Spectre。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_daneshmand.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_daneshmand.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">99. ARM MTE Performance in Practice</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Taehyun Noh (德克萨斯大学奥斯汀分校); Yingchen Wang (加州大学伯克利分校); Tal Garfinkel (谷歌); Mahesh Madhav (Ampere Computing); Daniel Moghimi (谷歌); Mattan Erez and Shravan Narayan (德克萨斯大学奥斯汀分校)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们对 ARM MTE 在四种不同微架构上的硬件性能进行了首次全面分析：Google Pixel 8 和 Pixel 9 上的 ARM Big（A7x）、Little（A5x）和 Performance（Cortex-X）核心，以及 Ampere Computing 的 AmpereOne CPU 核心。我们还包含了对 Apple M5 芯片上 MTE 的初步分析。我们在 MTE 的主要应用——概率性内存安全——上调查了其性能，涵盖 SPEC CPU 基准测试以及 RocksDB、Nginx、PostgreSQL 和 Memcached 等服务器工作负载。虽然 MTE 通常表现出适度的开销，但我们在某些基准测试上也看到了高达 6.64 倍的性能减速。我们识别了这些开销的微架构成因以及未来处理器可以在何处加以解决。随后，我们分析了 MTE 在更专业化安全应用中的性能，如内存追踪、检查时使用时（TOCTOU）防护、沙箱化和 CFI。在其中一些场景下，MTE 如今具有显著优势，而在其他场景下其收益微乎其微或取决于未来的硬件。最后，我们探讨了以往刻画 MTE 性能的工作在何处因方法论或实验误差而不完整或不正确。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_noh.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_noh.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h2 style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">100. Why Johnny Adopts Identity-Based Software Signing: A Usability Case Study of Sigstore</span></h2><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Kelechi G. Kalu, Sofia Okorafor, and Tanmay Singla (普渡大学); Sophie Chen (卡内基梅隆大学); Santiago Torres-Arias and James C. Davis (普渡大学)</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">软件签名是确保软件供应链中组件完整性和真实性的最稳健方法。传统签名工具给从业者带来了密钥管理和签名者身份识别的负担，造成了可用性挑战和安全风险。新一代签名工具已自动化了许多此类问题，但其可用性及其对实际采用和有效性的影响却知之甚少。可用性评估可以澄清新一代设计的成功程度，并突出改进的优先事项。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为填补这一空白，我们开展了对 Sigstore（新一代签名的先驱且广泛采用的典范）的首次可用性研究。通过对 17 位行业专家的访谈，我们考察了（1）从业者工具选择相关的问题和优势，（2）他们的签名工具使用方式及随时间演变的原因，以及（3）引发可用性问题的场景。我们的发现阐明了新一代签名工具的可用性因素，并为工具制作者、采用组织及研究界提供了建议。值得注意的是，新一代工具的不同组件展现出不同的成熟度和采用就绪程度，集成灵活性是一个常见的痛点，但可通过插件和 API 加以缓解。我们的结果将帮助新一代签名工具制作者进一步加强软件供应链安全。</span></p><p style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kalu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kalu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=96c42d16&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486124%26idx%3D1%26sn%3D5feac1df49b63f898b824fc3fd66ccd8">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 16 Jul 2026 20:22:00 +0800</pubDate>
    </item>
    <item>
      <title>USENIX Security 2026 — Cycle 1 论文清单与摘要（下）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486124&amp;idx=2&amp;sn=725a2ac1a1c3c5eb5898347a88ea60bc</link>
      <description></description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span>2026-07-16 20:22</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=8d8e8b9d&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sxqQV0Sn7E0Hm3u2MBfy8trQibtorfrvKdofQMdznGB3XMKk97WKZQLFticb06VLJX2d9NGl7OaVFyFHOBvlWw7IhXC1qNENsLu0%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="0"><span leaf="">101. Scribe: Low-memory SNARKs via Read-Write Streaming</span></h1><p data-layout-id="1" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Anubhav Baweja, Pratyush Mishra, Tushar Mopuri, Karan Newatia, and Steve Wang (宾夕法尼亚大学)</span></p><p data-layout-id="2" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="3" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">简洁非交互式知识论证（SNARK）使证明者能够为任意 NP 声明的有效性生成简短且可高效验证的证明。高效 SNARK 的近期构造激发了在广泛应用中使用它们的兴趣，但遗憾的是，在这些应用中部署 SNARK 面临一个关键瓶颈：即使是中等规模的声明，SNARK 证明者也需要大量的时间和内存来生成证明。尽管在减少证明者时间方面已有进展，但证明者内存仍是一个问题。</span></p><p data-layout-id="4" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们描述了 Scribe，一种新型低内存 SNARK，能够利用一种丰富但此前未被利用的资源——磁盘存储，即使在智能手机等廉价消费设备上也能高效地证明大规模声明。Scribe 的证明者不将其（大型）中间状态存储在 RAM 中，而是存储在磁盘上。为确保对状态的访问高效，我们在读写流式计算模型中设计了 Scribe 的证明者，使证明者只能以流式方式读取和修改其状态。</span></p><p data-layout-id="5" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们实现并评估了 Scribe 的证明者，结果表明，在商品硬件上，它可以轻松扩展到具有 228 个门的电路，同时使用不到 750MB 的内存，且与需要更多内存的最先进内存密集型基线（HyperPlonk [EUROCRYPT 2023]）相比，仅产生最小的证明延迟开销（10%）。</span></p><p data-layout-id="6" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baweja.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_baweja.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="8"><span leaf="">102. Identifying Provenance of Generative Text-to-Image Models</span></h1><p data-layout-id="9" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Anna Yoo Jeong Ha, Wenxin Ding, Stanley Wu, Shawn Shan, Haitao Zheng, and Ben Y. Zhao (芝加哥大学)</span></p><p data-layout-id="10" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="11" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">微调提供了一种快速且低成本的方式来生成新的文本到图像模型，这些模型往往与从头训练的模型难以区分。遗憾的是，对微调模型的虚假呈现给 AI 公司和用户都带来了问题，既抑制了竞争，又在模型质量及其训练过程的伦理方面误导了用户。</span></p><p data-layout-id="12" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出了一种模型溯源系统，仅需黑盒查询访问即可识别通过对现有文本到图像模型微调而生成的模型。我们的设计基于一项分析，即可以通过分析文本到图像模型对详细提示的响应来量化模型之间的特征空间差异。我们的系统分析模型输出，使用通用特征提取器提取视觉特征，并使用 Jensen-Shannon 散度将其分布与基础模型参考池的分布进行比较。随后应用统计假设检验来确定目标模型是从头训练还是微调而成，若为后者，则确定其可能的基础（父）模型。我们在七个广泛使用的扩散模型和众多微调变体上评估了该系统。结果表明，即使在图像后处理或权重扰动等对抗条件下，我们在模型谱系归因方面也具有很高的准确率。最后，我们通过追踪来自流行在线平台的野外模型的溯源，展示了系统在真实世界中的有效性。</span></p><p data-layout-id="13" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ha.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ha.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="15"><span leaf="">103. Semantics Over Syntax: Uncovering Pre-Authentication 5G Baseband Vulnerabilities</span></h1><p data-layout-id="16" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qiqing Huang and Xingyu Wang (布法罗大学); Wanda Guo and Guofei Gu (德克萨斯A&amp;M大学); Hongxin Hu (布法罗大学)</span></p><p data-layout-id="17" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="18" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代5G用户设备（UE）在建立认证与完整性保护之前，会在早期控制面交互过程中处理无线资源控制（RRC）配置消息。以往测试5G UE的工作大多集中于构造语法无效的输入。与此不同，我们证明语法有效但语义不一致的消息——即违反规范级字段约束或跨字段依赖关系的消息——能够将基带实现驱入非法状态，触发断言失败或调制解调器崩溃。这些发现揭示了认证前信令中的语义不一致性是5G UE实现中一个关键但尚未充分研究的攻击面。为弥补这一空白，我们提出约束引导的语义测试框架（Constraint-Guided Semantic Testing，CONSET），该框架系统性地抽取规范级约束，并利用这些约束生成针对性的语义违规以测试5G UE。CONSET将RRC消息解码为结构化字段，推导基于模式的规则，以证据有界的方式利用大语言模型（LLM）推断跨字段依赖关系，并生成语法有效但故意违反语义约束的测试用例。我们在商用与开源5G UE上对CONSET进行了评估。在商用智能手机上，通过负责任披露，它发现了7个此前未知的漏洞，其中包括3个高严重性CVE，影响64款芯片组型号及超过542款商用智能手机型号。在开源OAI UE上，CONSET额外触发了46个不同的崩溃点。</span></p><p data-layout-id="19" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-qiqing.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_huang-qiqing.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="21"><span leaf="">104. Window-based Membership Inference Attacks Against Fine-tuned Large Language Models</span></h1><p data-layout-id="22" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yuetian Chen, Yuntao Du, and Kaiyuan Zhang (普渡大学); Ashish Kundu (思科研究院); Charles Fleming (思科系统公司); Bruno Ribeiro and Ninghui Li (普渡大学)</span></p><p data-layout-id="23" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="24" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">针对大语言模型（LLM）的大多数成员推断攻击（MIA）依赖全局信号（如平均损失）来识别训练数据。然而，这种方法稀释了细微的、局部化的记忆信号，降低了攻击的有效性。我们对这种全局平均范式提出挑战，认为成员信号在局部上下文中更为显著。我们提出WBC（基于窗口的比较，Window-Based Comparison），通过滑动窗口结合基于符号的聚合来利用这一洞察。该方法在文本序列上滑动不同大小的窗口，每个窗口基于目标模型与参考模型之间的损失比较对成员身份进行二元投票。通过对几何级数间隔的多种窗口尺寸的投票进行集成，我们能够捕获从token级特征到短语级结构的记忆模式。在11个数据集上的大量实验表明，WBC显著优于现有基线方法，在低误报率阈值下取得更高的AUC分数，并将检测率提升2–3倍。我们的发现表明，聚合局部证据从根本上比全局平均更为有效，揭示了微调LLM中严重的隐私漏洞。</span></p><p data-layout-id="25" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-yuetian.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-yuetian.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="27"><span leaf="">105. Digital Risks and Coping Practices among Roblox Game Creators</span></h1><p data-layout-id="28" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qiurong Song, Rie Helene (Lindy) Hernandez, Xinning Gui, and Yubo Kou (宾夕法尼亚州立大学)</span></p><p data-layout-id="29" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="30" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">作为创作者经济日益增长的一部分，Roblox等游戏平台使数以百万计的用户能够设计、发布、推广并变现游戏。然而，在这些机遇之外，此类平台上的创作者也面临重大的安全、隐私与安保风险。尽管已有研究考察了社交媒体平台内容创作者面临的网络风险，但我们对游戏创作者的风险格局所知甚少。为弥补这一空白，我们访谈了20位Roblox创作者，以了解他们如何感知、经历并应对数字风险。我们的分析揭示了五类风险——平台、生产、组织、社区与技术——可能危及Roblox游戏创作者的情感、身体、人际关系及财务安全。我们还识别出诸如争取更公平报酬、寻求社区支持等应对策略。最后，我们提出了加强游戏创作者保护的建议。</span></p><p data-layout-id="31" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-qiurong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-qiurong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="33"><span leaf="">106. Paper Title Under Embargo</span></h1><p data-layout-id="34" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-style: italic;">（标题处于禁运状态，将在 USENIX Security 2026 会议开幕首日公开）</span></span></p><p data-layout-id="35" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Sam Crow (加州大学圣地亚哥分校); Stephen Checkoway (欧柏林学院); Patrick Mercier, Pat Pannuto, Stefan Savage, and Aaron Schulman (加州大学圣地亚哥分校)</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="39"><span leaf="">107. SoK: Attack and Defense Landscape of Agentic AI Systems</span></h1><p data-layout-id="40" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Juhee Kim (加州大学伯克利分校和首尔国立大学); Wenbo Guo (加州大学圣巴巴拉分校); Dawn Song (加州大学伯克利分校)</span></p><p data-layout-id="41" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="42" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">将大语言模型与非AI工具组件集成的AI智能体（AI agents）正快速涌现于现实应用中，提供了前所未有的自动化能力与灵活性。然而，这种灵活性引入了与传统软件系统不同的复杂安全挑战。在本文中，我们首次对AI智能体安全进行了全面的知识系统化梳理，分析了安全AI智能体系统的设计空间、攻击面与防御机制。此外，我们识别了这一新兴领域未来研究的开放性挑战。我们的工作为理解AI智能体安全风险与防御策略提供了首个系统性框架，可作为构建安全智能体系统并推进该关键领域研究的基础。</span></p><p data-layout-id="43" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-juhee-agentic.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kim-juhee-agentic.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="45"><span leaf="">108. Network-Level Prompt and Trait Leakage in Local Research Agents</span></h1><p data-layout-id="46" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hyejun Jeong, Mohammadreza Teymoorianfard, Abhinav Kumar, Amir Houmansadr, and Eugene Bagdasarian (马萨诸塞大学阿默斯特分校)</span></p><p data-layout-id="47" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="48" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们表明，Web与研究智能体（Web and Research Agents，WRA）——即基于语言模型、在互联网上调查复杂主题的系统——容易受到被动网络观察者的推断攻击。组织和个人出于隐私、法律或财务目的在本地部署WRA，使其暴露于DNS解析器、恶意ISP、VPN、Web代理以及企业或政府防火墙。然而，与人类偶发且稀疏的网页浏览不同，WRA对每个请求会访问70-140个域名，并具有独特的时序模式，从而产生独特的隐私风险。</span></p><p data-layout-id="49" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">具体而言，我们针对WRA演示了一种新型的提示词与用户特征泄露攻击，该攻击仅利用其网络级元数据（即访问的IP地址及其时序）。我们首先基于真实用户搜索查询和合成人格生成的查询，构建了一个新的WRA轨迹数据集。我们定义了一种行为度量指标（称为OBELS），以全面评估原始提示词与推断提示词之间的相似度，结果表明我们的攻击可恢复用户提示词中超过73%的功能与领域知识。扩展到多会话场景，我们以高准确率恢复了32个潜在特征中的最多19个。我们的攻击在部分可观测和含噪条件下仍然有效。最后，我们讨论了限制域名多样性或混淆轨迹的缓解策略，表明这些策略在效用影响可忽略的同时，将攻击有效性平均降低29%。</span></p><p data-layout-id="50" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jeong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jeong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="52"><span leaf="">109. NOIR: Privacy-Preserving Generation of Code with Open-Source LLMs</span></h1><p data-layout-id="53" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Khoa Nguyen (新泽西理工学院); That Khiem Ton (新泽西理工学院); NhatHai Phan (新泽西理工学院); Issa Khalil (哈马德·本·哈利法大学); Khang Tran and Cristian Borcea (新泽西理工学院); Ruoming Jin (肯特州立大学); Abdallah Khreishah (新泽西理工学院); My T. Thai (佛罗里达大学)</span></p><p data-layout-id="54" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="55" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">尽管大语言模型（LLM）驱动的代码生成能够提升软件开发效能，但由于服务提供商（云）能观察到客户端的提示词与生成代码，而在商业系统中这些内容可能属于专有资产，因此引入了知识产权与数据安全风险。为缓解这一问题，我们提出NOIR，这是首个保护客户端提示词与生成代码免受云端窥探的框架。NOIR在客户端使用编码器与解码器，对提示词的嵌入进行编码并发送至云端，从LLM获取增强后的嵌入，再在客户端本地解码生成代码。由于云端可能利用嵌入推断提示词与生成代码，NOIR引入了一种新机制来实现不可区分性——一种在token嵌入级别、针对提示词与代码所用词汇的本地差分隐私保护，并在客户端使用数据无关的随机化分词器。这些组件可有效防御诚实但好奇的云端发起的重建攻击与频率分析攻击。基于开源LLM的大量分析与结果表明，NOIR在多项基准上显著优于现有基线方法，包括Evalplus（MBPP与HumanEval，Pass@1分别为76.7和77.4）以及BigCodeBench（Pass@1为38.7，仅比原始LLM下降1.77%），同时在强隐私保护下抵御攻击。</span></p><p data-layout-id="56" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_nguyen.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_nguyen.pdf</a></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="57"><span leaf="">110. BADControl: Backdoor Attacks Against Control Systems</span></h1><p data-layout-id="58" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Luis Burbano (加州大学圣克鲁兹分校); Hampei Sasahara (东京科学大学); Ruoyu Song and Z. Berkay Celik (普渡大学); Alvaro A. Cardenas (加州大学圣克鲁兹分校)</span></p><p data-layout-id="59" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="60" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出BADCONTROL，这是首个使用物理触发器针对低级控制器的后门攻击。该攻击通过污染运行数据来植入漏洞，该漏洞可由来自环境的外部信号激活，例如自动驾驶应用中的特定驾驶操作或对抗性路面补丁。BADCONTROL通过使用投影梯度上升来修改数据，求解一个约束优化问题，使受控系统在目标频率处的频率响应最大化。该方法不同于针对深度学习（DL）与强化学习（RL）模型的后门攻击，后者操纵的是高维模型输入或奖励函数。我们还提出了两种防御方法：一种基于正则化，另一种基于鲁棒优化，用于限制触发器信号的最坏情况放大。这是通过一种专门的数学变换，将无限多种污染场景转化为单一可处理的优化问题来实现的。我们在比例-积分-微分（PID）控制器与线性二次型调节器（LQR）上通过仿真和物理实验对BADCONTROL进行了评估。在自适应巡航控制场景中，我们实现了100%的碰撞率；而在车道保持控制中，后门使受害车辆62%地驶入对向车道，而无后门时这两种情况均为0%。作为对比，面向自动驾驶车辆的最先进证伪框架在30次试验中仅识别出一次碰撞实例，凸显了本攻击的隐蔽性。</span></p><p data-layout-id="61" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_burbano.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_burbano.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="63"><span leaf="">111. Can we estimate privacy vulnerability of individual records? Towards Mitigating Attribute Inference Attacks on ML Models</span></h1><p data-layout-id="64" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Ehsanul Kabir and Najrin Sultana (宾夕法尼亚州立大学); Ninghui Li (普渡大学); Shagufta Mehnaz (宾夕法尼亚州立大学)</span></p><p data-layout-id="65" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="66" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">机器学习（ML）为各行各业带来了变革性应用，包括医疗保健、金融与客户分析等敏感领域。然而，ML模型容易发生隐私泄露，尤其是通过属性推断与模型反演攻击，这引发了对隐私关键领域数据保密性的担忧。现有防御所追求的目标远比专门防止属性推断攻击造成的隐私泄露更为宽泛，因而往往无法在不带来显著效用损失的情况下提供细粒度、感知脆弱性的保护。受此需求驱动，我们首先通过NeighVE——一种位于攻击方、旨在识别哪些个体记录更易受到推断的工具——研究记录级脆弱性估计。NeighVE揭示的洞察表明，记录级隐私泄露风险在很大程度上与模型架构和攻击策略无关，而是由数据集级特征决定，尤其是每条记录局部邻域内敏感属性的分布。基于这一洞察，我们提出VESL，一种受子空间学习启发的防御方法，可在将效用损失降至最低的同时缓解属性推断泄露。作为其平衡机制的副产品，VESL还改善了敏感属性间的公平性，并使NeighVE无法可靠地识别脆弱记录。作为辅助贡献，我们引入AttriVET，这是一种在多种场景下以超过90%的准确率预测哪些个体记录具有脆弱性的估计器，可支持感知风险的防御设计与审计。</span></p><p data-layout-id="67" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kabir.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_kabir.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="69"><span leaf="">112. Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution</span></h1><p data-layout-id="70" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chen Chen (南洋理工大学); Yuchen Sun and Jiaxin Gao (武汉大学); Xueluan Gong (南洋理工大学); Qian Wang (武汉大学); Ziyao Liu, Yongsen Zheng, and Kwok-Yan Lam (南洋理工大学)</span></p><p data-layout-id="71" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="72" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）取得了显著进展，在各类自然语言处理（NLP）任务中实现了卓越性能。然而，它们仍然容易受到后门攻击：在标准查询下模型表现正常，但当特定触发器被激活时会生成有害响应或非预期输出。现有后门防御在实践中要么缺乏全面性（仅关注狭窄的触发器设置、仅检测机制以及有限领域），要么无法抵御基于模型编辑、多触发器以及无触发器攻击等高级场景。在本文中，我们提出LETHE，一种通过知识稀释利用内部与外部机制消除LLM后门行为的新方法。在内部，LETHE利用一个轻量级数据集训练一个干净模型，随后将其与后门模型合并，通过在模型的参数化记忆中稀释后门影响来中和恶意行为。在外部，LETHE将与语义相关的良性证据融入提示词，以分散LLM对后门特征的注意力。在5个广泛使用的LLM上、跨分类与生成领域的实验结果表明，LETHE在抵御8种后门攻击时优于8种最先进的防御基线。LETHE将高级后门攻击的攻击成功率最高降低98%，同时保持模型效用。此外，LETHE已被证明高效且对自适应后门攻击具有鲁棒性。代码发布于<a href="https://github.com/Xxxxsir/Lethe。免责声明：本文包含可能具有冒犯性的内容。" target="_blank">https://github.com/Xxxxsir/Lethe。免责声明：本文包含可能具有冒犯性的内容。</a></span></p><p data-layout-id="73" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-chen.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-chen.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="75"><span leaf="">113. Distributed Synthesis of Differentially Private Tabular Datasets</span></h1><p data-layout-id="76" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yucheng Fu (弗吉尼亚大学); Tianyao Gu and Elaine Shi (卡内基梅隆大学); Tianhao Wang (弗吉尼亚大学)</span></p><p data-layout-id="77" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="78" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">差分隐私合成数据生成已成为一种在共享数据的同时保护个人隐私的强大工具。然而，当敏感数据的属性分布在多个实体（如医院、公司或政府机构）之间时，准确生成合成数据变得颇具挑战。尤其是，在未汇集整个私有数据集的情况下，难以捕获有信息量的统计相关性并利用其指导数据合成。为应对这一挑战，我们提出了一种面向分布式环境下差分隐私表格数据合成的安全多方计算协议。该协议包含两个新原语。第一个是利用分布式点函数高效估计垂直分布数据上二路边缘分布（属性的两两联合分布）的协议。第二个是通过在累积分布函数表中进行批量查找来生成噪声的协议。作为具体示范，我们构建了AIM（一种最先进的差分隐私数据合成算法）的分布式版本。我们的实现在达到与其集中式版本相同效用的同时，相比以往工作将端到端运行时间降低了数个数量级。例如，在真实广域网（WAN）环境中，我们能在24分钟内合成&#34;Adult&#34;数据集，而现有协议据估计需要57天。</span></p><p data-layout-id="79" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_fu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_fu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="81"><span leaf="">114. Trustworthy and Confidential SBOM Exchange</span></h1><p data-layout-id="82" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Eman Abu Ishgair and Chinenye Okafor (普渡大学); Marcela S. Melara (英特尔公司); Santiago Torres-Arias (普渡大学)</span></p><p data-layout-id="83" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="84" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">软件物料清单（SBOM）已成为一项监管要求，通过提供构成软件工件的组件的透明度来提升软件供应链安全与信任。然而，企业及受监管的软件供应商通常希望限制谁能查看其SBOM中记录的机密软件元数据，因为这些信息涉及知识产权或安全漏洞信息。为解决透明度与机密性之间的这一矛盾，我们提出Petra——一种SBOM交换系统，使软件供应商能够利用选择性加密，以可互操作的方式组合并分发经脱敏处理的SBOM数据。Petra使软件消费者能够在脱敏SBOM中搜索特定安全问题的答案，而不会泄露其未获授权访问的信息。Petra利用一种格式无关、防篡改的SBOM表示来生成高效且保护机密性的完整性证明，使相关方能对脱敏SBOM进行密码学审计并建立信任。在我们的Petra原型中，交换脱敏SBOM每份SBOM仅需不到1KB的额外开销，且SBOM解密在SBOM查询期间最多占1%的性能开销。</span></p><p data-layout-id="85" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ishgair.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ishgair.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="87"><span leaf="">115. Libra: Pattern-Scheduling Co-Optimization for Cross-Scheme FHE Code Generation over GPGPU</span></h1><p data-layout-id="88" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Song Bian, Yintai Sun, Zian Zhao, and Haowen Pan (北京航空航天大学); Mingzhe Zhang (无隶属单位); Zhenyu Guan (北京航空航天大学)</span></p><p data-layout-id="89" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="90" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出Libra，一种面向高并行计算架构、自动实现跨方案全同态加密（FHE）高效代码生成的编译器框架。虽然已知在单一应用中利用多种FHE方案可提升整体效率，但将跨方案FHE算子精确映射到通用图形处理器（GPGPU）等高性能架构上仍具挑战。为应对该挑战，Libra整合了FHE计算模式与硬件感知调度策略，构建了一个算法-硬件协同优化框架。具体而言，Libra为FHE定义了一种新颖的跨方案表示，抽象出每种FHE方案的通用程序模式。随后，我们基于从多种方案切换模式推导的FHE原语组合执行成本，对输出的FHE程序进行动态优化。接着，为加速GPU上的算子间执行，Libra引入一种将高层计算特征与低层执行计划相衔接的计算调度策略。通过所提出的模式-调度协同优化过程，Libra为GPGPU上的跨方案高精度FHE计算生成高效代码。实验结果表明，与最先进的跨方案工作相比，Libra在微基准上实现高达270倍加速，在应用上实现19倍加速，同时将计算单元和内存带宽利用率分别提升44%和36.1%。</span></p><p data-layout-id="91" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bian.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bian.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="93"><span leaf="">116. Efficient and High-Accuracy Secure Two-Party Protocols for a Class of Functions with Real-number Inputs</span></h1><p data-layout-id="94" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hao Guo and Zhaoqian Liu (香港中文大学（深圳）); Liqiang Peng (阿里巴巴集团); Shuaishuai Li (中关村实验室); Ximing Fu (香港中文大学（深圳）); Weiran Liu and Lin Qu (阿里巴巴集团)</span></p><p data-layout-id="95" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="96" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在两方秘密分享方案中，数值通常被编码为无符号整数uint(x)，而现实应用常常需要对带符号实数Real(x)进行计算。为实现实用函数的安全求值，必须能够从共享输入计算Real(x)，因为协议以份额作为输入。在USENIX&#39;25上，Guo等人提出了一种从份额高效计算带符号整数值int(x)的方法，可扩展用于计算Real(x)。然而，其方法对 x ∈ ZL 施加了严格的输入约束 |x| &lt; L⁄3，限制了其在现实场景中的适用性。在本工作中，我们将该约束显著放宽为对任意 B ≤ L⁄2 的 |x| &lt; B，其中 B = L⁄2 对应 x ∈ ZL 中的自然可表示范围。这放宽了限制，使得在宽松或无输入约束下计算Real(x)成为可能。在此基础之上，我们提出了一个通用框架，用于为一大类函数设计安全协议，包括整数除法（x⁄d的下取整）、三角函数（sin(x)）以及指数函数（e-x）。我们的实验评估表明，所提协议兼具高效率与高精度。值得注意的是，我们用于求值e-x的协议将通信开销降低至约SirNN（S&amp;P&#39;21）与Bolt（S&amp;P&#39;24）的31%，运行时分别加速达5.53倍和3.09倍。在精度方面，我们的协议最大ULP误差为1.435，而SirNN为2.64、Bolt为8.681。</span></p><p data-layout-id="97" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-hao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-hao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="99"><span leaf="">117. CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage</span></h1><p data-layout-id="100" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Na Li (南京理工大学网络空间安全学院（中国）); Yansong Gao (东南大学网络空间安全学院（中国）); Hongsheng Hu (上海交通大学计算机科学学院（中国）); Boyu Kuang (南京理工大学网络空间安全学院（中国）); Anmin Fu (南京理工大学网络空间安全学院（中国）以及南京理工大学计算机科学与工程学院（中国）)</span></p><p data-layout-id="101" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="102" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">模型压缩对于最小化深度学习（DL）模型的内存占用并加速推理至关重要。用户可根据自身资源与预算获取不同版本的压缩模型。然而，尽管现有压缩操作主要关注资源效率与模型性能之间的权衡，但压缩所引入的隐私风险却仍被忽视且未被充分理解。</span></p><p data-layout-id="103" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，针对典型的分类任务，我们从成员推断攻击（MIA）的视角提出CompLeak，这是首个隐私风险评估框架，考察了三种广泛使用的压缩配置——剪枝、量化和权重聚类——这三种配置均由谷歌TensorFlow-Lite（TF-Lite）商业模型压缩框架支持，前两种还由Facebook的PyTorch Mobile和微软NNI开源工具包支持。根据可获取的压缩模型数量和/或原始模型的可用性，CompLeak有三种变体。CompLeakNR首先采用现有MIA方法攻击每个单独的压缩模型，并发现不同压缩模型对成员与非成员的影响不同。当原始模型和一个压缩模型可用时，CompLeakSR将该压缩模型作为原始模型的参考，并通过结合两个模型的元信息（如置信度向量）揭示更多隐私。当多个压缩模型可用（无论是否可访问原始模型）时，CompLeakMR创新性地利用多个压缩版本的隐私泄露信息，显著放大整体隐私泄露。我们在六种多样的模型架构（从ResNet到BERT和GPT-2）以及五个图像与文本基准数据集上进行了大量实验。实验结果表明，CompLeakMR在包括0.1%误报率下的真阳性率（TPR）在内的所有评估指标上均取得最佳MIA性能，证明模型压缩加剧了隐私泄露。</span></p><p data-layout-id="104" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-na.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-na.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="106"><span leaf="">118. TAT: Attesting Trajectory Integrity of Industrial Robotic Arms</span></h1><p data-layout-id="107" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chengtao Yao, Chengcheng Zhao, Peng Cheng, and Jiming Chen (浙江大学)</span></p><p data-layout-id="108" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="109" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">工业机械臂是现代制造的核心，在关键领域得到广泛部署。运动是首要的安全关切，因为它是机械臂的基本能力，而对抗性操纵（例如篡改生产逻辑、定位或动力学）可能导致产品缺陷或物理损坏。远程证明是一种有前景的执行完整性验证机制。然而，现有方法侧重于控制流或数据流属性，未能捕获运动语义，限制了其充分验证机械臂物理执行的能力。</span></p><p data-layout-id="110" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出轨迹完整性（Trajectory Integrity，TI）作为一种新的安全属性，确保机械臂的运动符合其预期路径。为实施TI，我们设计了TAT，一个最小侵入式的证明框架，利用定时运动事件图（Timed Motion Event Graph）捕获运动语义，并结合事件测量与关节测量来验证实际运动。我们在开源机械臂平台上实现了TAT的软硬件原型。在真实任务程序上的评估表明，TAT至多带来2.30%的内存开销和0.14%的执行时间开销，证明了其性能与实用性。此外，我们在多种与运动相关的参数修改下评估了其证明能力，证实了其在轨迹完整性证明中的有效性。</span></p><p data-layout-id="111" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>暂未公开（论文处于禁运状态，PDF 将在 USENIX Security 2026 会议开幕首日发布）</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="113"><span leaf="">119. VidLeaks: Membership Inference Attacks Against Text-to-Video Models</span></h1><p data-layout-id="114" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Li Wang and Wenyu Chen (山东大学); Ning Yu (Eyeline Labs); Zheng Li and Shanqing Guo (山东大学)</span></p><p data-layout-id="115" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="116" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在海量网络规模数据集上训练的强大文本到视频（Text-to-Video，T2V）模型的激增，引发了关于版权与隐私侵犯的紧迫关切。成员推断攻击（MIA）为审计此类风险提供了一种规范化的工具，但现有技术针对图像或文本等静态数据设计，无法捕获视频生成的时空复杂性。尤其是，它们忽视了关键帧中记忆信号的稀疏性，以及随机时序动态所引入的不稳定性。</span></p><p data-layout-id="117" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们首次对针对T2V模型的MIA进行系统性研究，并引入一个新颖的框架VidLeaks，通过两种互补信号探测稀疏时序记忆：1）空间重建保真度（Spatial Reconstruction Fidelity，SRF），利用Top-K相似度从稀疏记忆的关键帧中放大空间记忆信号；2）时序生成稳定性（Temporal Generative Stability，TGS），通过测量多次查询间的语义一致性来捕获时序泄露。我们在三种逐步受限的黑盒设置下——有监督、基于参考和仅查询——对VidLeaks进行了评估。在三个代表性T2V模型上的实验揭示了严重的脆弱性：即便在最严格的仅查询设置下，VidLeaks在AnimateDiff上达到82.92%的AUC，在InstructVideo上达到97.01%的AUC，构成一种现实且可利用的隐私风险。我们的工作提供了首个具体证据，表明T2V模型会通过稀疏记忆与时序记忆泄露大量成员信息，为审计视频生成系统奠定了基础，并推动了新防御方法的开发。代码获取地址：<a href="https://zenodo.org/records/17972831。" target="_blank">https://zenodo.org/records/17972831。</a></span></p><p data-layout-id="118" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-li.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-li.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="120"><span leaf="">120. When Fun Turns Toxic: A First Look at Aggressive Advertising in Mini-games</span></h1><p data-layout-id="121" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Pei Chen, Geng Hong, Yicheng Qin, Huazhe Wang, Mengying Wu, and Min Yang (复旦大学); Ziru Zhao, Yuanpeng Zhu, and Tao Su (vivo移动通信有限公司)</span></p><p data-layout-id="122" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="123" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">小游戏已成为超级应用生态系统中的主导范式，使休闲游戏等轻量级服务能够瞬间触达数百万用户。虽然官方广告接口简化了变现流程，但集成的便捷性和监管的不足导致了激进且具有潜在欺骗性的广告行为，严重降低了用户体验。激进广告虽然不是恶意软件，但仍然通过滥用合法API来绕过审核、操纵用户交互并破坏平台信任，从而颠覆平台安全边界，构成系统性安全风险而非单纯的策略违规。</span></p><p data-layout-id="124" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们对小游戏中的激进广告进行了首次系统性安全分析。我们分析了九个小游戏平台的平台策略和开发者能力，并刻画了激进广告行为。我们进一步设计了一个可扩展的检测框架MAAD，并在三大主要平台（即微信、Facebook Instant Games和Quickgame）上进行了大规模测量，揭示了49.95%的小游戏存在激进广告，包括拥有超过10万用户评价的高人气作品。我们的分析进一步揭示了它们的破坏性行为模式，如游戏特定触发器、过度的弹窗频率和误导性策略，以及对抗性绕过技术。这些发现表明，激进广告构成了一种由当前执行机制结构性盲点所助长的广泛平台滥用形式。我们为加强平台治理、检测和长期生态韧性提供了可操作的建议。</span></p><p data-layout-id="125" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-pei.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chen-pei.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="127"><span leaf="">121. When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems</span></h1><p data-layout-id="128" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Shiqian Zhao (南洋理工大学); Jiayang Liu (南洋理工大学和东京科学大学（日本）); Yiming Li, Runyi Hu, and Xiaojun Jia (南洋理工大学); Wenshu Fan (电子科技大学); Xiaobao Wu and Xinfeng Li (南洋理工大学); Jie Zhang (新加坡科技研究局（A*STAR）CFAR与IHPC); Wei Dong and Tianwei Zhang (南洋理工大学); Luu Anh Tuan (南洋理工大学和VinUniversity)</span></p><p data-layout-id="129" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="130" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代文本到图像（T2I）生成系统（如DALL·E 3）利用记忆机制，在多轮交互中捕获关键信息以实现忠实的生成。尽管这一机制具有实用性，但对其安全分析远远滞后。在本文中，我们揭示了它可能加剧越狱攻击的风险。以往的攻击将不安全的目标提示融合为一个终极对抗提示，这很容易被检测到，或者由于解毒不足或过度而导致生成非不安全的图像。相比之下，我们提出在聊天会话开始时将恶意意图嵌入记忆中，从而解决上述局限性。</span></p><p data-layout-id="131" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">具体而言，我们提出了Inception，这是首个针对真实世界文本到图像生成系统的多轮越狱攻击，明确利用其记忆机制。Inception由两个关键模块组成：分割和递归。我们引入了Segmentation，一种保持语义的方法，可生成多轮提示。通过利用NLP分析技术，我们设计了策略，根据句子结构分解提示及其恶意意图，从而规避安全过滤器。递归进一步解决了无法通过简单分割分离的不安全子提示所带来的挑战。它首先扩展子提示，然后递归地调用分割。为便于多轮对抗提示的构建，我们构建了VisionFlow，一个集成了两阶段安全过滤器和工业级记忆机制的T2I仿真系统。实验结果表明，Inception成功诱导不安全图像的生成，在攻击成功率上以20.0%的优势超越了SOTA。我们还在真实的商业T2I生成平台上进行了实验，进一步验证了Inception在实际中的威胁。</span></p><p data-layout-id="132" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhao-shiqian.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhao-shiqian.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="134"><span leaf="">122. SoK: Security of Cyber-physical Systems Under Intentional Electromagnetic Interference Attacks</span></h1><p data-layout-id="135" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qinhong Jiang (香港理工大学); Yan Long (香港科技大学（广州）); Youqian Zhang (香港理工大学); Chen Yan and Xiaoyu Ji (浙江大学); Xiapu Luo (香港理工大学); Kevin Fu (东北大学); Jiannong Cao (香港理工大学); Wenyuan Xu (浙江大学)</span></p><p data-layout-id="136" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="137" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">故意电磁干扰（IEMI）攻击通过伪造计算机系统中的电信号——物理世界与数字世界之间的网关——已成为网络物理系统中日益普遍和具有破坏性的威胁，因其能够破坏或控制广泛的安全攸关和安全关键应用。现有的IEMI攻击研究通常高度针对特定设备，并利用分散且缺乏充分比较的攻击向量。缺乏基于模型的统一IEMI漏洞理解，既阻碍了可迁移的安全评估，也阻碍了面向可部署保护的有效跨学科合作。为弥补这一差距，本工作分析了80余个IEMI攻击与防御实例，提供了一个分析框架，建模对手如何实现IEMI耦合和样本操纵，以注入改变硬件行为并影响软件执行的恶意电磁能量。主要目标是推动该领域超越对脆弱实例的穷举式经验发现，转向适用于现有和未来网络物理系统的深入理论分析和主动防御策略。除了识别当前IEMI攻击与防御研究中的空白外，本工作还针对不同利益相关者群体的需求和角色，概述了未来工作的重要方向。为促进IEMI攻击的未来研究，我们在<a href="https://iemi-research-database.github.io/上发布并维护一个开源的IEMI研究数据库。" target="_blank">https://iemi-research-database.github.io/上发布并维护一个开源的IEMI研究数据库。</a></span></p><p data-layout-id="138" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jiang-qinhong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jiang-qinhong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="140"><span leaf="">123. Fend for Yourself! Backdoor Purification in Federated Graph Learning with an Evolving Knowledge Anchor</span></h1><p data-layout-id="141" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chengcheng Zhu and Yunlong Mao (南京大学); Jiale Zhang and Bosen Rao (扬州大学); Sheng Zhong (南京大学)</span></p><p data-layout-id="142" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="143" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">联邦图学习（FedGL）在保护隐私的同时实现了去中心化图数据的协作训练，但其分布式特性使其极易受到后门攻击。这些攻击通过注入恶意触发器来危害全局模型的完整性。然而，现有的防御方法在复杂图数据上往往无效，或依赖于可信服务器，与现代隐私保护技术产生架构冲突。为克服这些局限性，我们提出了GBHINDER，一种新颖且实用的无需可信服务器的防御框架，其中每个良性参与者自我防御。GBHINDER建立了一个良性循环：它利用自身可信的历史知识作为良性锚点来净化下载的全局模型，反过来，选择性地吸收全局模型的良性知识以逐步演化锚点本身。具体而言，这一循环由两个关键组件驱动。历史通道注意力正则化模块利用锚点约束全局模型的表示并破坏后门传播。为解决局部信任与全局协作之间的张力，自适应动量信息更新机制通过动态整合鲁棒的全局信息使锚点安全演化，确保锚点在联邦迭代中保持有效。在多个基准数据集上的大量实验表明，GBHINDER显著优于最先进（SOTA）的防御方法，成功将后门攻击成功率降至10%以下，同时在主任务上保持高准确率。</span></p><p data-layout-id="144" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-chengcheng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-chengcheng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="146"><span leaf="">124. InstantOMR: Oblivious Message Retrieval with Low Latency and Optimal Parallelizability</span></h1><p data-layout-id="147" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Haofei Liang (上海交通大学); Zeyu Liu (耶鲁大学); Eran Tromer (波士顿大学); Xiang Xie (Primus Labs); Yu Yu (上海交通大学)</span></p><p data-layout-id="148" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="149" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">不经意消息检索（OMR）解决了匿名消息系统和私有区块链中昂贵的消息检索过程。它使资源受限的接收方能够将消息的检测和检索外包，同时保护隐私。</span></p><p data-layout-id="150" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本工作提出了InstantOMR，一种新颖的OMR方案，在混合设计中将TFHE功能引导与标准RLWE操作相结合。InstantOMR专门针对低延迟和高并行性进行了优化。我们使用Primus-fhe库（以及基于TFHE-rs的估算）的实现表明，InstantOMR具有以下关键优势：</span></p><p data-layout-id="151" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="153"><span leaf="">125. Cracks in the Walled Garden: Dissecting the Gray-Market of Unauthorized iOS App Distribution via Ad Hoc Sideloading</span></h1><p data-layout-id="154" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yijing Liu, Yiming Zhang, Baojun Liu, and Haixin Duan (清华大学和BNRist)</span></p><p data-layout-id="155" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="156" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">Apple实施严格的代码签名，并要求通过其官方App Store分发应用。尽管如此，未授权应用仍通过侧载渠道传播。最初为开发者测试设计的Ad Hoc配置机制已成为其中一个渠道。它利用个人开发者证书和用户端签名，实现绕过Apple应用审核流程的未授权应用安装。随着时间的推移，这一做法已演变为一个结构化且普遍的灰色市场，连接了证书倒卖、第三方签名工具和未签名.ipa文件的分发。在本工作中，我们对这一市场进行了首次系统性研究，特别关注其在中国的一体化服务运营。通过以用户为中心的数据收集策略，我们识别了3,359个活跃的证书兑换签名站点，逆向分析了12款签名工具，并获取了8,216个分发的.ipa条目。我们的分析揭示了一个多层证书流通模型，转售利润率高达3,000%，并揭示了签名工具在代码签名中常用的技巧。大多数分发的应用是合法应用的修改版本，利用动态库注入来实现定制功能。此类修改破坏了应用和系统为用户提供的安全保护，使用户面临未授权操作、敏感数据外泄和系统能力利用等风险。总体而言，我们的发现揭示了一个成熟的灰色市场，它在公开运营的同时侵蚀iOS的信任模型，凸显了多方利益相关者进行针对性干预的必要性。</span></p><p data-layout-id="157" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-yijing.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-yijing.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="159"><span leaf="">126. ZipPIR: High-throughput Single-server PIR without Client-side Storage</span></h1><p data-layout-id="160" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Rasoul Akhavan Mahdavi, Abdulrahman Diaa, and Florian Kerschbaum (滑铁卢大学)</span></p><p data-layout-id="161" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="162" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">私有信息检索（PIR）允许客户端在不泄露访问哪个元素的情况下私密地访问数据库。基于Ring Learning with Errors（RLWE）的早期PIR协议证明了PIR的实用性，但吞吐量有限。另一种选择是，高吞吐量协议利用一个需要大量客户端存储的离线阶段（如SimplePIR中的提示），或在离线阶段产生高昂的通信成本（如Piano）。这些局限性与资源受限客户端的实际约束相冲突，并且在动态数据库中进一步加剧，因为更新需要昂贵的提示重新生成和重传。</span></p><p data-layout-id="163" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为应对这些挑战，我们提出了ZipPIR，一种高吞吐量PIR协议，将LWE密文压缩为显著更小的Paillier密文。ZipPIR利用离线阶段实现这种大小缩减，而不会在在线阶段产生相关计算成本。此外，在计算假设下，ZipPIR具有几乎静默的离线阶段，除初始公钥外不需要任何通信，使服务器能够在空闲时独立生成和更新提示，无需客户端交互。ZipPIR实现超过2 GB/s的吞吐量——可与SimplePIR等最先进协议相媲美——且无需客户端存储大型提示。对于1 GB数据库上的PIR，ZipPIR的吞吐量比无客户端存储的现有协议高出10倍，同时每客户端只需不到200 KB的服务器端存储，显著提升了实际部署的可扩展性。虽然此前使用Paillier的PIR协议效率很低，但ZipPIR是首个使用Paillier实现与最先进PIR协议相竞争吞吐量的PIR协议。</span></p><p data-layout-id="164" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_mahdavi.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_mahdavi.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="166"><span leaf="">127. Provable Secure Steganography Based on Adaptive Dynamic Sampling</span></h1><p data-layout-id="167" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Kaiyi Pang and Minhao Bai (清华大学)</span></p><p data-layout-id="168" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="169" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">由于广泛的监控，私密通信的安全日益面临风险。隐写术是一种将秘密消息嵌入无害载体中的技术，可在受监控的信道上实现隐蔽通信。可证明安全隐写术（PSS）确保正常模型输出与隐写输出之间的计算不可区分性，是该领域最先进的技术。然而，当前的PSS方法通常需要获取模型的显式分布。在本文中，我们提出了一种可证明安全的隐写方案，仅需一个接受种子作为输入的模型API。我们的核心机制涉及采样候选token集合并构建从可能的消息比特串到这些token的映射。通过将该映射应用于真实秘密消息来选择输出token，这可证明地保持了原始模型的分布。为确保正确解码，我们处理了多个候选消息映射到同一token的冲突情况，通过在有界大小范围内维护和策略性地扩展动态冲突集来实现。对三个真实世界数据集和三个大语言模型的广泛评估表明，我们基于采样的方法在效率和容量上与现有PSS方法相当。</span></p><p data-layout-id="170" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_pang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_pang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="172"><span leaf="">128. Arguzz: Testing zkVMs for Soundness and Completeness Bugs</span></h1><p data-layout-id="173" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Christoph Hochrainer (维也纳工业大学); Valentin Wüstholz (Diligence Security); Maria Christakis (维也纳工业大学)</span></p><p data-layout-id="174" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="175" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">零知识虚拟机（zkVM）越来越多地部署在去中心化应用和区块链rollup中，因为它们能够实现可验证的链下计算。这些VM执行通用程序（通常用Rust编写），并产生简洁的密码学证明。然而，zkVM非常复杂，其约束系统或执行逻辑中的漏洞可能导致严重的健全性（接受无效执行）或完备性（拒绝有效执行）问题。</span></p><p data-layout-id="176" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了ARGUZZ，首个用于测试zkVM健全性和完备性漏洞的自动化工具。为检测此类漏洞，ARGUZZ将蜕变测试的新颖变体与故障注入相结合。具体而言，它生成语义等价的程序对，将其合并为具有已知输出的单个Rust程序，并在zkVM中运行。通过向VM注入故障，ARGUZZ模拟恶意或有漏洞的证明者，以揭露过于薄弱的约束。</span></p><p data-layout-id="177" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们使用ARGUZZ测试了六个真实的zkVM——RISC Zero、Nexus、Jolt、SP1、OpenVM和Pico，并在其中三个中发现了11个漏洞。一个RISC Zero漏洞获得了50,000美元的赏金，尽管此前已有审计，这证明了对zkVM进行系统性测试的关键必要性。</span></p><p data-layout-id="178" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hochrainer.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hochrainer.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="180"><span leaf="">129. kSFS: Repurposing a Microkernel-like Interface for Fast and Secure In-Kernel Linux File Systems</span></h1><p data-layout-id="181" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Dinglan Peng and Pedro Fonseca (普渡大学)</span></p><p data-layout-id="182" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="183" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">文件系统被广泛使用且至关重要，但以复杂著称，是操作系统中漏洞的主要来源。最近的工作提出了引入内核内沙箱技术来隔离包括文件系统在内的内核组件。然而，一个定义良好且安全的边界——即所有不可信和可信内核组件之间的交互都应根据强威胁模型进行验证——往往被忽视。这种安全边界的缺乏尤其适用于Linux文件系统，它们依赖庞大且复杂的接口，并与VFS和块设备等许多内核子系统交互。定义这样的接口是沙箱化内核文件系统的具有挑战性的前提条件。</span></p><p data-layout-id="184" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们用kSFS解决了这一挑战，这是一个内核内沙箱文件系统框架。kSFS将FUSE协议（一种最初为Linux用户空间文件系统设计的类微内核接口）重新用作不可信沙箱内核文件系统的安全接口，具有强隔离保证。此外，kSFS将WebAssembly泛化到内核空间作为通用沙箱机制，并以最小的移植工作量实现与现有用户空间文件系统实现的兼容性。例如，使用kSFS将NTFS和exFAT实现从用户空间移植只需修改不到300行代码。在实现比Linux文件系统实现更好的安全性和可靠性的同时，kSFS实现了显著优于用户空间对应方案的性能。对于真实应用tar和RocksDB，kSFS的NTFS实现分别比用户空间基线性能高出29%和60倍，且仅比不安全的Linux实现性能低0%到52%。</span></p><p data-layout-id="185" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_peng-dinglan.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_peng-dinglan.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="187"><span leaf="">130. Sirens&#39; Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs</span></h1><p data-layout-id="188" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zijian Ling (华中科技大学和清华大学); Pingyi Hu, Xiuyong Gao, and Xiaojing Ma (华中科技大学); Man Zhou (华中科技大学和清华大学); Jun Feng and Songfeng Lu (华中科技大学); Dongmei Zhang and Bin Benjamin Zhu (微软公司)</span></p><p data-layout-id="189" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="190" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">语音驱动的大语言模型（LLM）越来越多地通过语音接口访问，通过开放的声学信道引入了新的安全风险。我们提出了Sirens&#39; Whisper（SWhisper），首个在现实黑盒条件下使用商用硬件对语音驱动LLM进行隐蔽基于提示攻击的实用框架。SWhisper能够在商用设备上实现任意目标基带音频（包括长而结构化的提示）的鲁棒、不可闻传输，方法是将其编码为近超声波形，在声学传输和麦克风非线性后忠实地解调。这通过一种简单而有效的方法来建模跨设备和环境的非线性信道特性，并结合轻量级信道反转预补偿来实现。在此高保真隐蔽信道的基础上，我们设计了一种语音感知的越狱生成方法，确保在语音驱动接口下的可懂度、简洁性和可迁移性。在商业和开源语音驱动LLM上的实验展示了强大的黑盒有效性。在商业模型上，SWhisper实现了高达0.94的不拒绝率（NR）和0.925的特定说服力（SC）。一项受控用户研究进一步表明，注入的越狱音频对于人类听者在感知上与纯背景播放无法区分。尽管越狱作为案例研究，但底层隐蔽声学信道使更广泛类别的高保真提示注入和命令执行攻击成为可能。</span></p><p data-layout-id="191" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ling.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_ling.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="193"><span leaf="">131. Autonomy Comes with Costs: Detecting Denial-of-Service Vulnerabilities Caused by Resource Abusing in LLM-based Agents</span></h1><p data-layout-id="194" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jiaqi Luo, Jiarun Dai, Fengyu Liu, Songyang Peng, Youkun Shi, Tong Bu, and Geng Hong (复旦大学); Xudong Pan (复旦大学和上海创新研究院); Yuan Zhang (复旦大学)</span></p><p data-layout-id="195" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="196" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于LLM的智能体最近引起了广泛关注。通过利用大语言模型（LLM）的语义理解能力，这些智能体可以根据用户请求自主执行复杂任务，如下载文件和总结内容。然而，缺乏全面的资源治理使其容易受到滥用，可能导致资源耗尽和拒绝服务（DoS）状态。</span></p><p data-layout-id="197" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本工作中，我们提出了对基于LLM智能体中资源管理的首次系统性安全研究。我们识别了资源生命周期管理的三种代表性模式，每种模式都为DoS利用提供了不同的途径。基于这些洞察，我们提出了AgentDoS，一种新颖的有向灰盒模糊测试框架，旨在检测由资源耗尽引起的DoS漏洞。AgentDoS首先分析智能体内的资源生命周期，然后利用LLM生成自然语言的功能特定种子提示，驱动智能体走向过度资源消耗。我们在20个广泛使用的开源基于LLM的智能体上评估了AgentDoS，发现了影响16个智能体的36个零日漏洞，其中15个在GitHub上拥有超过10,000颗星。截至目前，这些漏洞已获得15个CVE编号。</span></p><p data-layout-id="198" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_luo.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_luo.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="200"><span leaf="">132. Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems</span></h1><p data-layout-id="201" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Hongyan Chang, Ergute Bao, Xinjian Luo, and Ting Yu (穆罕默德·本·扎耶德人工智能大学)</span></p><p data-layout-id="202" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="203" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）越来越依赖从外部语料库中检索信息，这创造了新的攻击面：间接提示注入（IPI）。以往的研究强调了这一风险，但往往回避了最困难的步骤：确保恶意内容实际被检索到。在实践中，未经优化的IPI在自然查询下很少被检索到，这使其真实世界影响不明确。</span></p><p data-layout-id="204" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们通过将恶意内容分解为保证检索的触发片段和编码任意攻击目标的攻击片段来解决这一挑战。基于这一想法，我们设计了一种高效有效的黑盒攻击算法，构造紧凑的触发片段以保证任何攻击片段的检索。我们的攻击仅需对嵌入模型的API访问，成本低廉（在OpenAI的嵌入模型上每个目标用户查询低至0.21美元），并在11个基准和8个嵌入模型（包括开源模型和专有服务）上实现了近100%的检索率。</span></p><p data-layout-id="205" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于此攻击，我们展示了在自然查询和真实外部语料库下的首个端到端IPI利用，涵盖RAG和智能体系统，具有多样化的攻击目标。这些结果确立了IPI作为一种实用且严重的威胁：当用户发出自然查询以总结常见主题的电子邮件时，单封投毒电子邮件就足以迫使GPT-4o在多智能体工作流中以超过80%的成功率外泄SSH密钥。我们进一步评估了几种防御措施，发现它们不足以防止恶意文本的检索，凸显了检索作为一个关键的可利用漏洞。</span></p><p data-layout-id="206" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_chang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="208"><span leaf="">133. Inference Attacks Against Graph Generative Diffusion Models</span></h1><p data-layout-id="209" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Xiuling Wang and Xin Huang (香港浸会大学); Guibo Luo (北京大学); Jianliang Xu (香港浸会大学)</span></p><p data-layout-id="210" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="211" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">图生成扩散模型最近作为生成复杂图结构的强大范式而出现，有效捕获了图数据中错综复杂的依赖关系和关联。然而，与这些模型相关的隐私风险在很大程度上尚未被探索。在本文中，我们通过三类黑盒推理攻击来研究此类模型中的信息泄露。首先，我们设计了一种图重构攻击，能够从生成图中重构出与训练图结构相似的图。其次，我们提出了一种属性推理攻击，从生成图中推断训练图的属性，如平均图密度和密度分布。第三，我们开发了两种成员推理攻击，用于判断给定图是否存在于训练集中。在三种不同类型的图生成扩散模型和六个真实世界图上的大量实验证明了这些攻击的有效性，显著优于基线方法。最后，我们提出了两种防御机制来缓解这些推理攻击，并在防御强度和目标模型效用之间实现了比现有方法更好的权衡。我们的代码可在<a href="https://zenodo.org/records/17946102获取。" target="_blank">https://zenodo.org/records/17946102获取。</a></span></p><p data-layout-id="212" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-xiuling.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-xiuling.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="214"><span leaf="">134. Unlocking the True Potential of Decryption Failure Oracles: A Hybrid Adaptive-LDPC Attack on ML-KEM Using Imperfect Oracles</span></h1><p data-layout-id="215" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qian Guo, Denis Nabokov, and Thomas Johansson (隆德大学)</span></p><p data-layout-id="216" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="217" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">利用明文检查（PC）和解密失败（DF）预言机的侧信道攻击是对已部署后量子密码学的紧迫威胁。这些预言机可以从时序、功耗和微架构行为等有形泄漏源实例化，使其成为基于格、码和同源的主要方案的实际关注点。在本文中，我们重新审视了在ML-KEM上利用DF预言机的选择密文侧信道攻击。虽然DF预言机在基于格的方案中通常被认为不如其二进制PC对应物高效，但我们证明了其全部潜力在很大程度上尚未被实现。</span></p><p data-layout-id="218" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们引入了一种新颖的攻击框架，将自适应查询生成与低密度奇偶校验（LDPC）码的置信传播相结合。我们的方法在多个秘密系数上精心构造平衡的奇偶校验，最大化从每次预言机查询中提取的香农信息，即使在存在显著噪声的情况下也是如此。这种方法大幅减少了完整密钥恢复所需的查询次数，通过逼近理论香农信息界限实现了近最优效率。对于预言机准确率为95%的ML-KEM-768，我们的攻击仅需2950次查询（与香农下界的比率为1.35），证明了设计良好的DF攻击可以超越最先进二进制PC攻击的效率。</span></p><p data-layout-id="219" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为验证我们发现的实际影响，我们将该框架应用于近期的GoFetch攻击，展示了在这一真实世界微架构侧信道场景中的显著增益。我们的方法将所需测量迹减少了一个数量级以上，并消除了计算昂贵的后处理需求，使以前被认为难以处理的更高安全性方案上的完整密钥恢复成为可能。</span></p><p data-layout-id="220" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-qian.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_guo-qian.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="222"><span leaf="">135. vCause: Efficient and Verifiable Causality Analysis for Cloud-based Endpoint Auditing</span></h1><p data-layout-id="223" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qiyang Song, Qihang Zhou, Xiaoqi Jia, and Zhenyu Song (中国科学院信息工程研究所和中国科学院大学网络空间安全学院); Wenbo Jiang (电子科技大学); Heqing Huang (独立研究者); Yong Liu (奇安信科技集团股份有限公司); Dan Meng (中国科学院信息工程研究所和中国科学院大学网络空间安全学院)</span></p><p data-layout-id="224" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="225" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在云端终端审计中，安全管理员通常依赖云对日志衍生的版本化溯源图进行因果分析，以调查可疑攻击行为。然而，云可能不可信或被攻击者攻破，可能操纵最终的因果分析结果。因此，管理员可能无法准确理解攻击行为，从而无法实施有效的对策。这一风险凸显了确保因果分析完整性的防御方案的需求。虽然现有的防篡改日志方案和可信执行环境在这一任务上展现了前景，但它们并非专门为支持因果分析而设计，因此面临固有的安全性和效率局限性。</span></p><p data-layout-id="226" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了VCAUSE，一种用于云端终端审计的高效可验证因果分析系统。VCAUSE集成了两种认证数据结构：图累加器和可验证溯源图。这些数据结构能够验证因果分析中的两个关键步骤：（i）在版本化溯源图上查询兴趣节点，（ii）识别其因果相关组件。形式化安全分析和实验评估表明，VCAUSE能够实现安全可验证的因果分析，终端计算开销仅&lt;1%，云端为3.36%。</span></p><p data-layout-id="227" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-qiyang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_song-qiyang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="229"><span leaf="">136. CombiSan: Unifying Software Sanitizers for Comprehensive Fuzzing</span></h1><p data-layout-id="230" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Matteo Marini (罗马大学); Floris Gorter (阿姆斯特丹自由大学); Daniele Cono D&#39;Elia (罗马大学); Cristiano Giuffrida (阿姆斯特丹自由大学)</span></p><p data-layout-id="231" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="232" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代C/C++漏洞检测工作严重依赖模糊测试与软件消毒器的结合。然而，最流行的消毒器之间的互操作性有限。因此，开发者通常单独启用每个消毒器（如果启用的话），需要多次运行。这种顺序执行损害了性能，并以非统一的方式测试代码。</span></p><p data-layout-id="233" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在本文中，我们提出了CombiSan，一种模糊测试优化的消毒器，可同时检测三种最流行消毒器（ASan、MSan和UBSan）所覆盖的所有可寻址性、未初始化内存和其他未定义行为问题。CombiSan采用统一的影子内存设计，高效跟踪程序内存每个字节的可寻址性和初始化状态。此外，CombiSan的插桩与其他未定义行为类别的最先进检测无缝集成。由于不同消毒器发现的漏洞可能因提前终止执行而相互掩盖，CombiSan将所有聚合问题的分析推迟到测试用例完成时进行。</span></p><p data-layout-id="234" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在我们的评估中，CombiSan在OSS-Fuzz每日测试的10个程序中检测到81个新漏洞。平均而言，使用CombiSan的模糊测试比顺序测试ASan+UBSan和MSan快1.7倍。此外，我们的结果表明，尽管运行时间显著减少，CombiSan具有与这些消毒器相同的漏洞检测准确率。</span></p><p data-layout-id="235" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_marini.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_marini.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="237"><span leaf="">137. KernelRCA: Facilitating Root Cause Analysis of Memory Corruptions in Linux Kernel with Contextual Causality Chain</span></h1><p data-layout-id="238" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Kangzheng Gu, Yifan Zhang, Yuan Zhang, and Min Yang (复旦大学)</span></p><p data-layout-id="239" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="240" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">持续模糊测试基础设施已经发现了大量漏洞。在此背景下，自动根因分析（RCA）被提出以减少理解漏洞根因所需的高昂人工成本。然而，现有的根因表示采用孤立形式设计。分析人员仍需手动推断包括调用上下文和数据依赖在内的完整漏洞触发流程，而由于操作系统的复杂性，这对操作系统内核而言极为困难。</span></p><p data-layout-id="241" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了上下文因果链（contextual causality chain，CC-chain），一种新颖的根因表示方法，可直观反映 Linux 内核中内存损坏的完整漏洞触发流程。CC-chain 展示了促成漏洞的指令，以解释导致漏洞的相应异常行为，同时呈现这些指令之间的调用上下文和数据依赖，帮助分析人员快速理解漏洞的发生机制。为自动构建 CC-chain，我们设计了根因分析系统 KernelRCA，包括选择性追踪、上下文信息恢复和链式根因分析。KernelRCA 成功诊断了 Linux 内核中 54 种真实世界的内存损坏，表现优于现有的崩溃报告和 KASAN 报告。一项用户研究表明，KernelRCA 的报告显著提升了人工分析人员对漏洞的理解和修复效率。</span></p><p data-layout-id="242" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gu-kangzheng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_gu-kangzheng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="244"><span leaf="">138. Sy-FAR: Symmetry-based Fair Adversarial Robustness</span></h1><p data-layout-id="245" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Haneen Najjar, Eyal Ronen, and Mahmood Sharif (特拉维夫大学)</span></p><p data-layout-id="246" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="247" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">安全关键型机器学习（ML）系统，如人脸识别系统，容易受到对抗样本的攻击，包括现实世界中物理可实现的攻击。已有多种方法被提出来增强 ML 的对抗鲁棒性；然而，这些方法通常会引发不公平的鲁棒性：从某些类别（如个人）或群体（如性别）发起攻击往往比从其他类别或群体发起攻击更容易。已有若干技术被开发用于在寻求类别间完美公平性的同时提升对抗鲁棒性。然而，先前的工作主要集中在安全性和公平性不太关键的场景（例如对汽车和船只等物体的分类）。</span></p><p data-layout-id="248" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的洞察是，在现实世界中公平性关键的任务（如人脸识别）中实现完美的对等性通常是不可行的——某些类别（如兄弟姐妹）可能高度相似，导致它们之间出现更多误分类。相反，我们认为寻求对称性——即从类别 i 到 j 的攻击成功率与从 j 到 i 的攻击成功率相同——更为可行。直观上，对称性是可取的，因为在大多数领域中类别相似性是一种对称关系。此外，正如我们从理论上证明的那样，个体之间的对称性会诱导任意子群体集合之间的对称性，这与群体公平性常常难以实现的其他公平性概念形成对比。</span></p><p data-layout-id="249" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们开发了 Sy-FAR，这是一种在同时优化对抗鲁棒性的前提下鼓励对称性的技术，并使用五个数据集、三种模型架构对其进行了广泛评估，包括针对目标式和非目标式现实攻击的评估。结果表明，与最先进的方法相比，Sy-FAR 显著提升了公平的对抗鲁棒性。此外，我们发现 Sy-FAR 在多次运行中速度更快、一致性更好。值得注意的是，Sy-FAR 还缓解了我们在本工作中发现的另一种不公平性——在引入对称性后，对抗样本最可能被分类进入的目标类别变得明显更不容易受攻击。</span></p><p data-layout-id="250" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_najjar.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_najjar.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="252"><span leaf="">139. Ajax: Fast Threshold Fully Homomorphic Encryption without Noise Flooding</span></h1><p data-layout-id="253" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhenkai Hu (上海交通大学 and 密码学国家重点实验室); Haofei Liang (上海交通大学); Xiao Wang (西北大学); Xiang Xie (华东师范大学 and Primus Lab); Kang Yang (密码学国家重点实验室); Yu Yu (上海交通大学); Wenhao Zhang (西北大学)</span></p><p data-layout-id="254" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="255" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">门限全同态加密（ThFHE）使多方能够在加密数据上执行任意计算，同时密钥分布在各方之间。设计 ThFHE 的主要任务是为 FHE 方案构建门限密钥生成和解密协议。在现有 FHE 方案中，类 FHEW 密码系统具有快速自举和参数小的优势。然而，已知的 ThFHE 方案使用“噪声淹没”（noise-flooding）技术来实现门限解密，这要求要么使用大参数，要么通过自举切换到具有大参数的方案，导致解密过程缓慢。此外，在密钥生成方面，现有的 ThFHE 方案要么假设通用 MPC 或可信设置，要么产生与参与方数量 n 成线性关系的噪声增长。</span></p><p data-layout-id="256" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了快速 ThFHE 方案 Ajax，通过为类 FHEW 密码系统设计门限密钥生成和解密协议来实现。具体而言，对于门限解密，我们消除了噪声淹没的需求，转而提出一种基于不同环上随机双重共享的新技术——“先掩码后公开”（mask-then-open），同时保持了参数小的优势。对于门限密钥生成，我们展示了一种简单的方法，在诚实多数设定下（至多 t=(n-1)/2 方被腐败），将噪声增长从 n 倍降低到 max(0.038n,2) 倍。我们的端到端实现报告了在 1 Gbps 带宽和 1 ms 延迟的网络下，对于 n=3（分别为 n=21）方，生成一组密钥和解密单个密文的运行时间分别为 17.6 s 和 0.9 ms（分别为 91.9 s 和 4.4 ms）。与最先进的实现相比，我们的协议在 t=1 到 t=13 的不同网络延迟下，将门限解密协议的端到端性能提升了至少 5.7× 至 283.6×。我们的方法也可应用于 BGV、BFV 和 CKKS 等其他类型的 FHE 方案。</span></p><p data-layout-id="257" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_hu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="259"><span leaf="">140. FirmReBugger: A Benchmark Framework for Monolithic Firmware Fuzzers</span></h1><p data-layout-id="260" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Mathew Duong, Michael Chesser, and Guy Farrelly (阿德莱德大学); Surya Nepal (澳大利亚联邦科学与工业研究组织 Data61); Damith C. Ranasinghe (阿德莱德大学)</span></p><p data-layout-id="261" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="262" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">单片固件无处不在。因此，固件模糊测试是一个活跃的研究领域，不断有新进展来解决该领域的独特挑战。然而，通过推导代码覆盖率和独特崩溃等指标来理解和评估改进效果存在问题，这引发了对可靠的基于漏洞的基准测试的需求。为满足这一需求，我们设计并构建了 FirmReBugger，一个利用真实的、多样的、基于漏洞的基准来公平评估单片固件模糊测试工具的整体框架。</span></p><p data-layout-id="263" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">FirmReBugger 提出使用漏洞预言机——即漏洞描述符的 C 语法表达式——配合解释器来自动化分析并准确报告发现的漏洞，区分已检测、已触发、已到达和未到达等状态。重要的是，我们的基准测试理念不修改目标二进制文件，仅通过重放模糊测试种子来将基准测试实现与模糊测试工具隔离，同时提供一种简单的方式来扩展新的漏洞预言机。</span></p><p data-layout-id="264" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">此外，通过分析模糊测试障碍，我们创建了 FirmBench，一组包含 313 个软件漏洞预言机的多样真实世界二进制目标。结合我们对单片固件模糊测试障碍的分析，该基准为快速评估未来的进展提供了支持。我们将 FirmReBugger 实现为一个 FuzzBench-for-Firmware 类型的服务，并使用 FirmBench 以可复现研究的方式评估了 9 个最先进的单片固件模糊测试工具，耗费 10 CPU 年的计算量来报告我们的发现。</span></p><p data-layout-id="265" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_duong.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_duong.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="267"><span leaf="">141. Differential Trust: Dynamic Multi-Authority Anonymous Credentials with Epoch-Weighted Updates</span></h1><p data-layout-id="268" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chen Li (天津大学); Jianting Ning (浙江理工大学); Xiulong Liu (天津大学); Yulin Liu (武汉大学)</span></p><p data-layout-id="269" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="270" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">匿名凭证（AC）是隐私保护认证的基础，允许用户证明其属性拥有权而不泄露身份。最先进的 AC 将凭证签发分散到多个机构，通常采用 Shamir 秘密共享或聚合签名等技术。虽然这种方法增强了系统鲁棒性并消除了单点故障，但它在凭证签发阶段对所有机构一视同仁。这种统一处理忽视了不同机构所持有的不同信任度或权益。这一限制在现代去中心化系统（如 Proof-of-Stake 网络）中尤为突出，因为节点间固有的信任差异化无法在凭证签发过程中被利用。</span></p><p data-layout-id="271" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为解决这一限制，我们提出了基于纪元权重的多机构匿名凭证（Multi-Authority Anonymous Credentials with Epoch-Based Weights，MA-ACEW）概念，这是第一个在凭证签发中考虑机构权重分布的多机构匿名凭证（MA-AC）模型。关键的是，MA-ACEW 在机构权重分布跨纪元变化时能够高效更新凭证。MA-ACEW 的核心是我们新颖的纪元绑定 Pointcheval-Sanders 签名（Epoch-Bound Pointcheval-Sanders Signature，EB-PS）原语，它将签名绑定到特定时间纪元。这种时间绑定既支持纪元内基于权重的凭证签发，又支持跨纪元的高效非交互式凭证更新。我们形式化了 EB-PS 的 EUF-eCMA 不可伪造性要求，并证明在新的 STB-GPS 假设下我们的构造满足该要求。随后我们证明 MA-ACEW 构造实现了不可伪造性、匿名性和盲性。最后，我们展示了基准测试结果，证明了 EB-PS 和 MA-ACEW 的效率。值得注意的是，展示一个由 128 个部分凭证聚合而成的凭证平均仅需 10.68 ms。</span></p><p data-layout-id="272" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-chen.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-chen.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="274"><span leaf="">142. From Texts to Rules: Generating Sigma Rules with Large Language Models from Cyber Threat Reports</span></h1><p data-layout-id="275" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yongxin Cai (广州大学); Jing Qiu (广州大学 and 鹏城实验室); Qingming Li (浙江大学); Du Cheng (清华大学); Lei Chen (香港科技大学)</span></p><p data-layout-id="276" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="277" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">网络威胁报告（CTR）提供了安全系统检测规则所需的可操作情报。大语言模型（LLM）可以通过其解析和生成能力，作为 CTR 到规则翻译的桥梁。然而，CTR 中的高层抽象与规则中的底层机器语义之间的语义脱节和领域特定约束，从根本上阻碍了检测规则的准确生成。</span></p><p data-layout-id="278" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文证明了 CTR 中的 shell 命令可以有效地转换为安全系统的 Sigma 检测规则。为此，我们提出了 SIGMERGE，一个端到端框架，通过构建语义中间层作为桥梁，从 CTR 文本生成 Sigma 规则。SIGMERGE 框架按语义层次由高到低分层组织三个模块：（1）信息提取模块，高层级，利用多子序列算法和微调的领域专用 LLM，实现准确的 MITRE ATT&amp;CK 战术、技术和程序（TTP）及命令提取；（2）攻击描述生成模块，中间层级，采用偏好优化调优和闭环自验证来缓解语义脱节；（3）Sigma 规则生成模块，机器级，利用参数优化检索算法来解决领域特定约束。我们构建了 7 个数据集用于训练并进行了广泛实验。为验证 SIGMERGE，我们使用 23 个指标对其与 16 个基线和 13 个 LLM 进行了评估，并开展了 10 个与真实安全系统集成的案例研究，以展示其有效性和效率。此外，SIGMERGE 已向官方仓库贡献了 4 条新颖的 Sigma 规则，所有规则均已被正式接受。</span></p><p data-layout-id="279" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cai.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cai.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="281"><span leaf="">143. Static Detection of TOCTOU Bugs Caused by Kernel Races</span></h1><p data-layout-id="282" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Gui-Dong Han, Jia-Ju Bai, Qiu-Ji Chen, and Jiqiang Lu (北京航空航天大学)</span></p><p data-layout-id="283" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="284" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">TOCTOU（检查时间到使用时间）漏洞是内核代码中一个众所周知的安全问题，因为它绕过安全检查并导致可能引发系统崩溃和权限提升等严重问题的异常行为。根据我们对 Linux 内核补丁的研究，内核竞争是内核 TOCTOU 漏洞最常见的根因。然而，由于内核并发逻辑的复杂性和线程调度的非确定性，目前尚无专注于检测由内核竞争引起的 TOCTOU 漏洞的系统性方法。</span></p><p data-layout-id="285" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文设计了 KERAT，这是第一个用于检测由内核竞争引起的 TOCTOU 漏洞的系统性静态分析方法。实际上，此类 TOCTOU 漏洞是由特定共享变量的检查-使用操作的原子性违例引入的。因此，KERAT 通过从内核代码中静态挖掘和检查共享变量的原子性规则来执行漏洞检测。具体而言，KERAT 具有两个关键技术：（1）原子性规则挖掘方法，用于有效识别哪把锁应该保护哪个共享变量的检查-使用操作；（2）基于状态的验证策略，利用常见漏洞模式的状态机编码来检测违反已挖掘原子性规则的 TOCTOU 漏洞。我们在 Linux-6.8 和 FreeBSD-14.1 上评估了 KERAT，发现了 351 个真实漏洞。其中 287 个被确认为有害，65 个已被内核开发者确认。10 个漏洞已获得 CVE 编号。</span></p><p data-layout-id="286" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_han.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_han.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="288"><span leaf="">144. Paper Title Under Embargo</span></h1><p data-layout-id="289" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-style: italic;">（标题处于禁运状态，将在 USENIX Security 2026 会议开幕首日公开）</span></span></p><p data-layout-id="290" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Benedict Schlüter, Christoph Wech, and Shweta Shinde (苏黎世联邦理工学院)</span></p><p data-layout-id="291" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span>（本论文摘要目前处于禁运状态，将在 USENIX Security 2026 会议开幕首日公开发布。）</span></p><p data-layout-id="292" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>暂未公开（论文处于禁运状态，PDF 将在 USENIX Security 2026 会议开幕首日发布）</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="294"><span leaf="">145. OS-Sanitizer: System-wide Latent Defect Inference in Linux Applications</span></h1><p data-layout-id="295" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Addison Crump, Sahil Sihag, Florian Bauckholt, and Keno Hassler (德国亥姆霍兹信息安全研究中心（CISPA）); Thorsten Holz (马克斯·普朗克安全与隐私研究所)</span></p><p data-layout-id="296" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="297" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">动态测试在历史上一直专注于发现软件产生不良行为的情形，通常通过触发故障或不良状态来实现。然而，此类测试通常局限于通过示例来发现这些场景。我们能否通过检查良性行为来判断软件可能产生不良行为？本文通过利用 eBPF 在 Linux 应用中实现动态缺陷推断来探索这一问题。eBPF 作为一种系统自省工具具有独特优势，它从用户态和内核态事件中收集数据，并作为程序在内核中进行处理。我们的原型 OS-Sanitizer 使用启发式方法实现了此类 eBPF 程序，能够在整个系统的所有应用中报告疑似存在的缺陷。从概念上讲，OS-Sanitizer 将静态测试中的代码异味（code smells）理念引入动态测试，同时受益于运行时事件的洞察。通过这种方式，我们推断出软件中潜在上下文缺陷的存在，这些缺陷仅在某些环境中才会引发故障，或者以其他方式难以测试。我们从性能、复杂性、可维护性和可用性的角度考虑并评估了这种方法的优缺点，区分了 eBPF 的理论极限与我们原型的具体限制。针对众所周知的软件缺陷类型，我们在广泛使用的应用中发现了超过 40 个问题（包括严重漏洞），其中一些已存在超过十年并存在于大多数 Linux 发行版上。我们的发现表明，动态缺陷推断既可行又有效，凸显了在软件测试中扩展这一探索不足方向的机会。</span></p><p data-layout-id="298" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_crump.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_crump.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="300"><span leaf="">146. SafeFFI: Efficient Sanitization at the Boundary Between Safe and Unsafe Code in Rust and Mixed-Language Applications</span></h1><p data-layout-id="301" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Oliver Braunsdorf and Tim Lange (慕尼黑大学); Konrad Hohentanner and Julian Horsch (弗劳恩霍夫 AISEC); Johannes Kinder (慕尼黑大学)</span></p><p data-layout-id="302" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="303" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">非安全 Rust 代码对于与 C/C++ 库的互操作性和实现底层数据结构是必要的，但它可能在原本内存安全的 Rust 程序中引发内存安全违例。消毒器可以在运行时捕获此类内存错误，但即便对于 Rust 类型系统保证安全的内存访问，也会引入许多不必要的检查。我们引入了 SafeFFI，一个用于优化 Rust 二进制文件中内存安全插桩的系统，使检查发生在非安全代码和安全代码的边界处，将内存安全的执行从消毒器移交给 Rust 类型系统。与先前的方法不同，我们的设计避免了昂贵的全程序分析；因此，它产生的编译时开销显著更低（2.01× 对比超过 5.91×）。在一组流行的 Rust crate 上，SafeFFI 将消毒器检查减少了多达 79.63%，同时仍能检测出我们已知漏洞 Rust 代码数据集中的所有内存安全违例。</span></p><p data-layout-id="304" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_braunsdorf.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_braunsdorf.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="306"><span leaf="">147. GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs</span></h1><p data-layout-id="307" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Lichao Wu, Sasha Behrouzi, and Mohamadreza Rostami (达姆施塔特工业大学); Stjepan Picek (萨格勒布大学 and 拉德堡德大学); Ahmad-Reza Sadeghi (达姆施塔特工业大学)</span></p><p data-layout-id="308" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="309" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">混合专家（Mixture-of-Experts，MoE）架构通过每个输入仅激活稀疏参数子集，推进了大语言模型（LLM）的扩展，在降低计算成本的同时实现了最先进的性能。随着这些模型越来越多地部署在关键领域，理解和加强其对齐机制对于防止有害输出至关重要。然而，现有的 LLM 安全研究几乎完全集中于密集架构，对 MoE 的独特安全特性在很大程度上尚未进行考察。MoE 的模块化、稀疏激活设计表明，安全机制的运作方式可能与密集模型不同，这引发了对其鲁棒性的质疑。</span></p><p data-layout-id="310" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了 GateBreaker，这是首个无需训练、轻量级且架构无关的攻击框架，能够在推理时破坏现代 MoE LLM 的安全对齐。GateBreaker 分三个阶段运行：（i）门控级画像，识别在有害输入上被不成比例路由的安全专家；（ii）专家级定位，在安全专家内部定位安全结构；（iii）定向安全移除，禁用已识别的安全结构以破坏安全对齐。我们的研究表明，MoE 安全性集中在由稀疏路由协调的一小部分神经元中。选择性地禁用这些神经元——最多占目标专家层的 2.9% 的神经元——在效用损失有限的情况下，将针对八个最新对齐 MoE LLM 的平均攻击成功率（ASR）从 7.4% 提升至 64.9%。这些安全神经元可在同系列模型间迁移，通过单次迁移攻击将 ASR 从 17.9% 提升至 67.7%。此外，GateBreaker 可泛化到五个 MoE 视觉语言模型（VLM），在不安全图像输入上达到 60.9% 的 ASR。据我们所知，此前没有工作能在对抗 MoE LLM 时达到如此程度的有效性。</span></p><p data-layout-id="311" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wu.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wu.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="313"><span leaf="">148. Bridges to Self: Silent Web-to-App Tracking on Mobile via Localhost</span></h1><p data-layout-id="314" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Tim Vlummens (COSIC, 鲁汶大学); Aniketh Girish and Nipuna Weerasekara (IMDEA Networks Institute); Frederik Zuiderveen Borgesius and Gunes Acar (拉德堡德大学); Narseo Vallina-Rodriguez (IMDEA Networks Institute)</span></p><p data-layout-id="315" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="316" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">现代浏览器和移动操作系统利用沙箱和进程隔离来分离 Web 和 App 上下文。然而，本文表明这些隔离保证在实践中已经被——而且曾被——Meta 和 Yandex 在 Android 设备上打破，以实现将 Web 追踪与原生身份关联起来的跨上下文追踪。</span></p><p data-layout-id="317" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">结合来自美国和欧盟视角的大规模 Web 爬虫和系统性 Android App 分析，我们刻画了一类此前未被记录的 Web-to-App 追踪范式，该范式利用 HTTP(S)、WebSocket 和 WebRTC 等 Web 标准，通过 localhost 连接移动和 Web 上下文。通过将匿名 Web Cookie 与长期存在的原生用户 ID 关联，这些通道实现了持久且隐蔽的跨上下文追踪和去匿名化。这种新技术能够击败 Cookie 清除、隐身模式、移动广告标识符（MAID）重置、VPN 以及 Android 的工作/个人配置文件分离等保护措施。我们进一步表明，Meta Pixel 和 Yandex Metrica 在接受 Cookie 同意横幅之前就启动了 localhost 桥接。我们评估了浏览器在响应我们负责任披露后对这些攻击的修补工作和防御措施，以及即将推出的本地网络访问（Local Network Access，LNA）权限，该权限引入了访问 localhost 和本地网络地址的用户提示。在此过程中，我们发现了绕过这些保护的额外侧信道，利用（i）WebRTC 中的全局单播 IPv6 地址；以及（ii）对 *.local 域名的 mDNS 查询。我们的结果连同所附的法律分析，揭示了结构性缺陷，以及重新审视平台和浏览器的隔离原则、威胁与信任模型、协议标准和应用审核流程的必要性，以防止未来的跨上下文滥用。</span></p><p data-layout-id="318" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_vlummens.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_vlummens.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="320"><span leaf="">149. WILD Attack: Stealthy Undermining of Wi-Fi-Based Geolocation Through Remote Crowdsourced Data Injection</span></h1><p data-layout-id="321" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Changjia Zhu, Xiao Han, Parush Gera, Zhuo Lu, Tempestt Neal, and Yao Liu (南佛罗里达大学)</span></p><p data-layout-id="322" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="323" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">传统的 Wi-Fi 定位系统（WPS）欺骗攻击虽然看似有效，但由于缺乏隐蔽性和持久性，未能引起重大的 WPS 安全关注。本文提出了一种新颖的 WILD 攻击，通过颠覆 WPS 的核心基础设施——位置查询表（Location Lookup Table，LLT）——来破坏 WPS 安全。在此攻击中，攻击者远程提交针对目标 Wi-Fi 接入点的伪造众包报告，诱导 WPS 提供商基于伪造而非合法数据更新 LLT。我们考察了四个广泛部署的 WPS 提供商——Google、Apple、A-Map 和 WiGLE——发现它们都接受伪造报告，并采用不同的策略来解决合法数据与伪造数据之间的冲突。利用这些策略，攻击者可以诱导两种形式的 LLT 颠覆：LLT 条目篡改和 LLT 条目移除，两者即使在攻击者停止活动后仍可持续数周。我们进一步展示了三个案例研究来说明 WILD 攻击的现实影响，并提出了缓解此类威胁的对策。</span></p><p data-layout-id="324" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-changjia.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-changjia.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="326"><span leaf="">150. When Updates Backfire: A Black-Box Security Analysis of Desktop Software Update Mechanisms</span></h1><p data-layout-id="327" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jie Wan, Pengcheng Xia, and Haoyu Wang (华中科技大学)</span></p><p data-layout-id="328" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="329" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">桌面软件已成为现代计算的重要组成部分，而软件更新是修补漏洞和提供安全修复的主要机制。然而，更新过程本身引入了新的攻击面，特别是当更新数据的验证不完整或执行不当时。尽管如此，桌面更新中的这些风险较少被研究，因为大多数更新客户端是闭源且复杂的。我们提出了 UpdSight，一个通过在真实环境中模拟中间人（MitM）攻击来测试更新安全性的黑盒框架。UpdSight 通过模拟中间人场景、在更新过程中拦截流量，并自动验证关键弱点是否存在来运作。通过结合流量拦截、载荷完整性检查和行为监控，UpdSight 对多种软件类别的更新信任模型提供了全面评估。我们将 UpdSight 应用于 85 个广泛使用的桌面应用。结果显示了 22 个可利用漏洞，包括降级攻击、清单篡改、安装程序劫持和路径遍历。其中 16 个已被供应商确认。此外，还分配了 5 个 CVE 标识符，覆盖了 8 个软件产品中的漏洞。我们的发现揭示了反复出现的设计缺陷，如未签名的清单和弱回滚检查，这些缺陷使攻击者能够通过更新通道获得代码执行能力。</span></p><p data-layout-id="330" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wan.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wan.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="332"><span leaf="">151. TopFeaRe: Locating Critical State of Adversarial Resilience for Graphs Regarding Topology-Feature Entanglement</span></h1><p data-layout-id="333" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Xinxin Fan (人工智能安全国家重点实验室, 中国科学院计算技术研究所; and 中国科学院大学); Wenxiong Chen (大连理工大学; and 人工智能安全国家重点实验室, 中国科学院计算技术研究所); Quanliang Jing (中国科学院计算技术研究所); Chi Lin (大连理工大学); Shaoye Luo (人工智能安全国家重点实验室, 中国科学院计算技术研究所; and 中国科学院大学); Wenbo Song (大连理工大学; and 人工智能安全国家重点实验室, 中国科学院计算技术研究所); Yunfeng Lu (北京航空航天大学)</span></p><p data-layout-id="334" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="335" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">图对抗攻击通常从拓扑/结构和节点特征两个视角产生，两者都代表了当今深度学习模型所学到的至关重要的特征。尽管目前已有一些防御对策被提出，但它们未能揭示这两个方面为何必要的内在原因以及它们如何被充分融合以协同学习图表示。针对这一问题，本文借助复杂动力系统（CDS）学科中的平衡点理论，提出了一种通过定位图的对抗韧性临界状态来实现对抗防御的方法。简言之，本工作有三项创新：i）对抗攻击建模，即将图体系映射到 CDS，利用动力系统的振荡来建模对抗扰动行为；ii）针对扰动图的二维拓扑-特征纠缠函数设计，即将图拓扑和节点特征投影为两个特征空间，定义二维纠缠扰动函数来表示对抗攻击下的动态方差；iii）对抗韧性临界状态定位，即利用平衡点理论，借助扰动映射的二维函数来定位图的攻击韧性临界状态。最后，在五个常用真实数据集上的多方面实验验证了所提方法的有效性，结果表明我们的方法在四种代表性图对抗攻击下能显著优于最先进的基线方法。</span></p><p data-layout-id="336" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_fan.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_fan.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="338"><span leaf="">152. BatchBoot: Fast Batched Bootstrapping for TFHE scheme and Practical Applications</span></h1><p data-layout-id="339" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhihao Li (蚂蚁数字科技, 蚂蚁集团); Hongyu Wang (山西大学); Yuan Zhao and Lichun Li (蚂蚁数字科技, 蚂蚁集团); Zhiwei Wang (网络空间安全防御国家重点实验室, 中国科学院信息工程研究所); Jiaxing He, Changzheng Wei, and Ying Yan (蚂蚁数字科技, 蚂蚁集团); Lifeng Guo (山西大学)</span></p><p data-layout-id="340" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="341" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于环面的全同态加密（TFHE）以其独特的自举机制为特征，该机制在执行任意计算的同时刷新噪声预算。然而，该机制表现出有限的可扩展性，因为它一次只能处理单个加密消息。为解决这一问题，近期研究提出了批量自举方案，使 TFHE 能够并行处理密文，从而实现有前景的摊销效益。尽管取得了这些进展，这一新兴方向仍未被充分探索，留下了充足的研究空间。</span></p><p data-layout-id="342" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文提出了 BatchBoot，一个高效的 TFHE 批量自举框架，实现了加密消息的摊销处理。具体而言，本工作有三项关键贡献。首先，我们重新设计了核心子模块，即同态多项式乘法，以大幅减少对昂贵的 FFT 操作的依赖。其次，我们提出了一种稀疏感知的消息打包策略，灵活支持不同的打包规模。第三，我们将函数自举扩展到电路自举，从而大大增强了所支持函数的表达能力。这些贡献共同使 BatchBoot 相比最先进的批量方案（Guimarães et al., CCS&#39;25）实现了 2.4× 的加速，相比非批量 TFHE-rs 实现实现了 43.8× 的提升。</span></p><p data-layout-id="343" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在应用层面，我们通过两个实际用例展示了 BatchBoot 的通用性。首先，我们提出了非均衡设定下首个基于 TFHE 的 PSI 协议，相比最优的基于 BFV 的方案（PEPSI, USENIX Security&#39;24），实现了 294× 的通信成本降低和 4.1× 的加速。其次，我们基于 BatchCBoot 设计了一个 8 位 FHE 指令集，相比现有结果（Wang et al., CCS&#39;25）实现了高达 5.4× 的加速。</span></p><p data-layout-id="344" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-zhihao.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-zhihao.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="346"><span leaf="">153. The Adverse Effects of Omitting Records in Differential Privacy: How Sampling and Suppression Degrade the Privacy–Utility Tradeoff</span></h1><p data-layout-id="347" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Àlex Miranda-Pascual (卡尔斯鲁厄理工学院 and 加泰罗尼亚理工大学); Javier Parra-Arnau (加泰罗尼亚理工大学); Thorsten Strufe (卡尔斯鲁厄理工学院)</span></p><p data-layout-id="348" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="349" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">采样以在差分隐私（DP）中的隐私放大效应而闻名，并常被认为可以通过允许噪声减少来提升 DP 机制的效用。本文进一步表明，后一假设是有缺陷的：在相同隐私水平下衡量效用时，作为预处理的采样在所有经典 DP 机制——Laplace、Gaussian、exponential 和 report noisy max——中，以及采样的近期应用（如聚类）中，由于省略记录导致的效用损失而持续产生惩罚。</span></p><p data-layout-id="350" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在此分析的基础上，我们将抑制作为选择或省略记录的广义方法进行研究。通过发展对该技术的理论分析，我们在无界近似 DP 下推导出了任意抑制策略的隐私界。我们发现，所测试的抑制策略同样未能改善隐私-效用权衡。令人惊讶的是，均匀采样作为最佳抑制方法之一脱颖而出——尽管其仍具有降级效应。我们的结果对 DP 实践中常见的预处理假设提出了质疑。</span></p><p data-layout-id="351" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_miranda-pascual.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_miranda-pascual.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="353"><span leaf="">154. VeCT: Secure and Efficient Constant-Time Code Rewriting with Vector Extensions</span></h1><p data-layout-id="354" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Qisheng Jiang and Danfeng Zhang (杜克大学)</span></p><p data-layout-id="355" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="356" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">时序侧信道允许攻击者通过分析受害程序的执行时间来提取秘密。常时间（CT）编程范式通过数据流/控制流线性化（DFL/CFL）来防御时序攻击。然而，改写后的常时间代码通常会显著增加原始代码的内存占用，导致可观的开销。我们提出了 VeCT，一种基于编译器的代码改写器，利用向量扩展在保持常时间保证的同时提升性能。我们首先应用严格的统计测试，为实现细节保密的 AVX-512 指令推导出实用的“安全使用”规则；该分析还揭示了一个现有最先进常时间改写器中此前未知的漏洞。在这些规则的指导下，VeCT 引入了一种新颖的策略，消除改写代码中不必要的数据加载，并启用向量化以进一步提高效率。我们基于 LLVM 实现了 VeCT，可将代码自动转换为基于 AVX-512 的常时间等价代码。在 AES 和 Blowfish 等实际应用上，VeCT 将转换后代码的开销较现有最先进方法降低了最多 98.9%，同时保持了常时间行为。</span></p><p data-layout-id="357" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jiang-qisheng.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_jiang-qisheng.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="359"><span leaf="">155. Memclave: Secure In-memory Enclave for Untrusted Hosts</span></h1><p data-layout-id="360" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Amit Choudhari (德国亥姆霍兹信息安全研究中心（CISPA）); Fabian van Rissenbeck (多特蒙德工业大学); Christian Rossow (德国亥姆霍兹信息安全研究中心（CISPA）)</span></p><p data-layout-id="361" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="362" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">云平台在多租户环境下运行数据密集型工作负载，频繁的 CPU-内存流量可能通过缓存侧信道泄漏访问模式。存内处理（PIM）设备（如 UPMEM）将计算迁移至 DRAM 中，大幅减少数据移动并缩小 CPU 缓存占用。然而，商用 PIM 架构暴露了由主机编程的控制平面和主机共享的模块内存，使得设备端驻留的代码和数据易受被攻陷主机的威胁。现有的安全 PIM 方案要么添加加密/访问控制硬件，要么依赖重量级的主机端加密协议，给实际部署带来复杂性。</span></p><p data-layout-id="363" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们提出了 Memclave，一个纯软件框架，无需修改硬件即可为商用 PIM 带来代码完整性和数据机密性。经过 TPM 认证的管理程序在启动时将 PIM 的控制平面与主机访问永久隔离。在每个存内核心上，可信加载器对用户内核进行认证，并建立每会话受保护的数据通路。Memclave 保留了原有的编程模型和内核代码：主机应用将少量数据移动调用替换为安全的插入式替代，使得可信计算基保持小巧，移植成本低。我们在现成 UPMEM DIMM 上实现了 Memclave，并在 PrIM 基准测试套件上进行了评估，涵盖异构的内存访问、计算和同步模式。在一次性 100ms 认证加载后，存内内核时间接近 PIM 基线：多层感知机（MLP）在实用规模下保持在 1.5× 以内，广度优先搜索（BFS）在部分图上为 1.1×，随着前沿层级数的增加仅有适度上升。</span></p><p data-layout-id="364" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_choudhari.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_choudhari.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="366"><span leaf="">156. Cracking Federated Privacy: Initialization-Resilient Gradient Inversion with Fine-Grained Reconstruction</span></h1><p data-layout-id="367" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Kaiming Zhu, Jinsheng Yang, Siyang Guo, Huaqian Qin, Taiyu Wang, Junbo Wang, Yuhong Nan, and Zibin Zheng (中山大学)</span></p><p data-layout-id="368" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="369" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">联邦学习（FL）仍易受梯度反演攻击（GIA）的威胁，共享梯度可能泄露客户端的私有数据。现有攻击在早期初始化变化下表现不佳，且往往产生粗糙的重建结果。本文中，我们识别出共享梯度中的稀疏性变化是这种敏感性的主要来源，并提出一种对初始化具有鲁棒性的 GIA，采用由粗到细的设计，实现细粒度恢复。粗粒度阶段对齐梯度方向并约束非零项以缓解稀疏性变化，细粒度阶段通过结合余弦距离与变形曼哈顿距离项的混合度量来优化幅度对齐。针对五个基线的大量实验表明，在 CIFAR-10/100 上敏感初始化条件下 PSNR 提升高达 200%（25.4 → 47.7 dB），并在四个数据集和整个 FL 生命周期中持续实现细粒度恢复。我们的方法在各批大小和本地步数上保持与 SOTA 基线相当的竞争力，并揭示了若干流行模型上持续存在的隐私泄漏以及现有防御的不足，凸显了加强隐私保护机制的迫切需求。</span></p><p data-layout-id="370" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-kaiming.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_zhu-kaiming.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="372"><span leaf="">157. Estimating the Amount of Script-generated Traffic in a Mixture</span></h1><p data-layout-id="373" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Cormac Herley (微软研究院)</span></p><p data-layout-id="374" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="375" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们研究在混合流量中估计机器人生成流量占比的问题。即，当接收到的流量为 α · Clean + (1-α) · Bot 时，我们寻求估计 α。该问题主要针对流量试图伪装为人类生成的情况（例如点击欺诈、虚假社交媒体互动等）。</span></p><p data-layout-id="376" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">当干净流量中至少有一对特征相互独立时（例如地理分布的时间不变性），我们证明获取 α 的上界等价于寻找使一个简单目标函数最大化的秩一矩阵。我们给出了一种高效的求解方法，并推导了该界的紧致性。当数据有限时，误差分析极为重要，因为秩一矩阵的采样版本并不会精确为秩一。我们为估计量推导了置信区间，使我们能够确信所得到的是真正的上界。</span></p><p data-layout-id="377" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们通过实证验证了上述发现。首先，分别使用随机秩一矩阵和满秩矩阵作为干净流量和机器人流量的分布，通过蒙特卡洛模拟验证了准确性。其次，我们考察了 Twitter（现为 X）的数据。在公开市场上出售的拥有大量粉丝的 Twitter 账户被标记为拥有 &gt;90% 的机器人粉丝，而若干学术会议和知名研究人员的账户则被标记为 &lt;20%。我们在任意干净/机器人组成的 Twitter 账户群体上验证了准确性。</span></p><p data-layout-id="378" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>源页面未提供 PDF 链接</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="380"><span leaf="">158. Quorus: Efficient, Scalable Threshold ML-DSA Signatures from MPC</span></h1><p data-layout-id="381" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Alexander Bienstock, Leo de Castro, Daniel Escudero, Antigoni Polychroniadou, and Akira Takahashi (J.P. Morgan AlgoCRYPT CoE and J.P. Morgan AI Research)</span></p><p data-layout-id="382" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="383" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">阈值签名协议将秘密签名密钥划分给多个参与方，使得任何超过阈值的子集都能共同生成签名。后量子（PQ）阈值签名正受到广泛研究，尤其是在 NIST 发布阈值方案征集之后，但大多数解决方案专注于专门设计的、对阈值友好的签名方案。然而，分布式证书颁发机构和数字货币等实际应用要求签名可在现有标准化流程下验证。随着 NIST 对 PQ 签名的标准化以及产业界的持续部署，设计一种与 NIST 标准化验证兼容的高效阈值方案仍是一个关键挑战。</span></p><p data-layout-id="384" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本文中，我们提出了首个高效且可扩展的模块格数字签名算法（ML-DSA）多方生成解决方案，ML-DSA 是 NIST 的 PQ 签名标准之一。我们的贡献有两方面。首先，我们提出了一种适用于高效多方计算（MPC）的 ML-DSA 签名算法变体，并证明该变体与原始 ML-DSA 方案具有相同的安全性。其次，我们提出了若干高效且可扩展的 MPC 协议来实例化阈值签名功能。我们的协议在每次拒绝采样轮次中仅需每方 150 KB 的在线通信即可生成阈值签名。此外，我们在诚实多数设定下实例化了这些协议，从而避免任何额外的公钥假设。</span></p><p data-layout-id="385" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们的签名在所有安全级别下可通过与 ML-DSA 相同的实现进行验证，签名和验证密钥大小与 ML-DSA 一致；此前的基于格的阈值方案无法同时匹配这两项大小。我们的解决方案提供了首个与 NIST 标准化验证兼容、可扩展至任意参与方数量且无需新假设的阈值后量子签名生成方法。</span></p><p data-layout-id="386" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bienstock.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_bienstock.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="388"><span leaf="">159. Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and Profit</span></h1><p data-layout-id="389" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Alejandro Cuevas (卡内基梅隆大学); Manoel Horta Ribeiro (普林斯顿大学); Nicolas Christin (卡内基梅隆大学)</span></p><p data-layout-id="390" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="391" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在线内容创作者花费大量时间和精力，通过一个漫长且往往艰辛的过程来建立自己的用户群，需要找到合适的“细分领域”来服务受众。那么，一位以猫咪表情包闻名的成熟内容创作为何要彻底重塑其频道，转而推广加密货币服务或报道选举新闻事件？如果他们真的这样做了，其现有订阅者难道不会注意到吗？</span></p><p data-layout-id="392" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们探讨了频道重新利用的问题，即频道更改其身份和内容。我们首先刻画了一个“二手”社交媒体账户市场，在为期 6 个月的观察期内，其销售额超过 100 万美元。观察这 6 个月内被（转）售的 YouTube 频道，我们发现相当数量（53%）的频道被用于传播政策敏感内容，且往往未受到任何处罚。更令人惊讶的是，这些频道似乎在增加而非减少订阅者。</span></p><p data-layout-id="393" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我们使用从生态有效的代理中采样的两组共 140 万个 YouTube 账户快照，估计了频道重新利用在“野外”的普遍程度。在 3 个月期间，我们估计有 0.25% 的频道被重新利用。通过一系列实验，我们确认这些被重新利用的频道与被售频道共享若干特征——主要是它们具有显著高比例的政策敏感内容。在被重新利用的频道中，我们发现了类似于影响力操纵行动中所使用的频道，以及用于金融诈骗的频道。被重新利用的频道拥有庞大的受众；在两组观测样本中，被重新利用的频道合计分别拥有 1.93 亿和 4400 万订阅者。我们认为，购买现有受众和已建立账户所附带的信誉，对受经济和意识形态驱动的攻击者是有利的。这一现象并非 YouTube 独有，我们推断培育有机受众的市场将会增长，尤其是在缺乏技术或其他方面干预的情况下。</span></p><p data-layout-id="394" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cuevas.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_cuevas.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="396"><span leaf="">160. Khost: KVM-based Near Native MCU Firmware Rehosting</span></h1><p data-layout-id="397" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Chunlin Wang, Yicheng Yang, Yuan Zhang, Haoyu Xiao, Yifan Zhang, and Jiarun Dai (复旦大学)</span></p><p data-layout-id="398" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="399" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于微控制器单元（MCU）的设备构成了物联网基础设施的关键层，因此确保其安全至关重要。基于重新托管（rehosting）的动态 MCU 固件分析是保障这些设备安全的有效方法。然而，现有的重新托管框架由于模拟而普遍存在显著的性能开销，或执行范围受限。</span></p><p data-layout-id="400" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为解决这些局限性，我们提出了 Khost，一个接近原生速度、保留执行范围的重新托管框架。它通过引入轻量级扩展 CPU、辅助页表和软件中断控制器来扩展 KVM，使 MCU 固件能够以最低开销在高性能平台上重新托管。它还提供了内存映射 I/O（MMIO）监控器以实现快速外设交互，以及一个固件封装器以支持覆盖收集并灵活配置现有模糊测试引擎。在两个标准基准上的评估表明，与 QEMU 相比，Khost 将复杂计算任务的开销降低了 90.0% 至 95.5%，将 MCU 系统级操作的开销降低了最多 98.5%。此外，在 12 个真实固件上使用 Khost 进行模糊测试实现了高达 197.5× 的吞吐量提升，并将基本块覆盖率较现有模糊测试工具提高了 6 倍。此外，Khost 成功发现了 5 个此前未知的漏洞。</span></p><p data-layout-id="401" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-chunlin.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-chunlin.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="403"><span leaf="">161. Attesting Model Lineage by Consisted Knowledge Evolution with Fine-Tuning Trajectory</span></h1><p data-layout-id="404" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Zhuoyi Shang, Jiasen Li, and Pengzhen Chen (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院; and 网络空间安全防御重点实验室); Yanwei Liu (中国科学院信息工程研究所; and 网络空间安全防御重点实验室); Xiaoyan Gu (中国科学院信息工程研究所; 中国科学院大学网络空间安全学院; and 网络空间安全防御重点实验室); Weiping Wang (中国科学院信息工程研究所)</span></p><p data-layout-id="405" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="406" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">深度学习中的微调技术在模型之间催生了新兴的谱系关系。这种谱系为解决未经授权的模型再分发和虚假声称模型来源等安全问题提供了一个有前景的视角，这些问题在缺乏稳健谱系验证机制的开源权重模型库中尤为紧迫。现有的模型谱系检测方法主要依赖于静态架构相似性，不足以捕捉真正谱系关系所基于的知识动态演化。受人类进化的遗传机制启发，我们通过验证知识演化与参数修改的联合轨迹来解决模型谱系认证问题。为此，我们提出了一种新颖的模型谱系认证框架。在我们的框架中，首先利用模型编辑来量化微调引入的参数级变化。随后，我们引入一种新颖的知识向量化机制，在探针样本的辅助下，将编辑后模型中演化的知识提炼为紧凑表示。探针策略针对不同类型的模型族进行了适配。这些嵌入作为验证跨模型知识关系算术一致性的基础，从而实现对模型谱系的稳健认证。大量实验评估表明了我们的方法在多种真实世界对抗场景中的有效性和鲁棒性。我们的方法在包括分类器、扩散模型和大语言模型在内的广泛模型类型上持续实现可靠的谱系验证。</span></p><p data-layout-id="407" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shang.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_shang.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="409"><span leaf="">162. SONIC: Concurrent Oblivious RAM &amp; Data Structures for Low-Latency and High-Throughput</span></h1><p data-layout-id="410" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Nihal Talur and Ioannis Demertzis (加州大学圣克鲁兹分校)</span></p><p data-layout-id="411" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="412" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">仅依赖加密进行隐私保护计算容易受到泄漏滥用/访问模式攻击。TEE 虽然成本效益高，但也易受侧信道攻击。不经意原语，如不经意内存（ORAM）和不经意数据结构（ODS），是缓解这些风险的有效构建模块，可隐藏内存访问模式和侧信道信息。应用场景涵盖私密联系人发现（Signal）、匿名密钥透明性、加密邮件搜索、加密/不经意数据库、匿名通信（Sparta/SP&#39;25）、隐私联邦学习、LLM 隐私（Compass/OSDI&#39;25）以及更广泛的机密计算。</span></p><p data-layout-id="413" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">基于树的 ORAM（EnigMap（USENIX&#39;23）、GraphOS（PVLDB&#39;23）、Oblix（SP&#39;18））提供低延迟但并行度有限，即使对小数据集也难以突破 1K req/s 的吞吐量。基于分区的方案如 Snoopy（SOSP&#39;21）将数据划分到多个子 ORAM 中，每个子 ORAM 通过由传入请求构建的不经意哈希表并行扫描其分片，以大量牺牲延迟换取高吞吐——理论上实现线性可扩展性。实践中，Snoopy 的性能取决于每个子 ORAM 能多快完成其顺序扫描以不超过延迟目标——这限制了服务器利用率和吞吐量。虽然 TB 级数据集在理论上可通过增加服务器实现，但实际中需要 1000+ 台服务器。</span></p><p data-layout-id="414" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本工作中，我们调和了上述低延迟与高吞吐 ORAM 方案之间的割裂局面。我们引入 SONIC：首个面向硬件飞地的 ORAM，用 RingORAM 替换了 PathORAM（EnigMap、GraphOS 和 Oblix 所使用）。我们的设计是首个实现最低 150K req/s 且最高可达 2M req/s 吞吐量（单服务器）的低延迟 ORAM，攻克了所有树 ORAM 构造（包括 RingORAM）的核心挑战，如克服顺序驱逐瓶颈、实现高效批量驱逐，以及提供无锁的访问、重排和暂存操作。SONIC 的 ORAM 访问吞吐量比开源 EnigMap 实现高 28-197×，比 GraphOS 高 158-1065×（N=227，块大小 64 字节）。SONIC 提供了多种利用 ORAM 并发性来构建不经意数据结构（如 OMAP）的方法。最后，SONIC 可作为 Snoopy 子 ORAM 的直接替代品以提供更实用的可扩展性——1TB 现在仅需 32 台服务器即可处理，而非数千台。</span></p><p data-layout-id="415" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span>暂未公开（论文处于禁运状态，PDF 将在 USENIX Security 2026 会议开幕首日发布）</span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="417"><span leaf="">163. E2E-AKMA: An End-to-End Secure and Privacy-Enhancing AKMA Protocol Against the Anchor Function Compromise</span></h1><p data-layout-id="418" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Yueming Li (中国科学院软件研究所; and 中国科学院大学); Long Chen (中国科学院软件研究所; and 中国科学院系统软件重点实验室); Qianwen Gao (中国科学院软件研究所; and 中国科学院大学); Zhenfeng Zhang (中国科学院软件研究所; and 中国科学院系统软件重点实验室)</span></p><p data-layout-id="419" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="420" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">应用认证与密钥管理（AKMA）系统是 3GPP 制定的一种近期协议，预计将成为 5G 标准的关键组成部分。AKMA 使应用服务提供商能够将用户认证流程委托给移动网络运营商，从而免去这些提供商自行存储和管理认证相关数据的需求。这种委托提升了认证流程的效率，但同时也引入了某些值得深入分析和缓解的安全与隐私挑战。</span></p><p data-layout-id="421" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">5G AKMA 服务由 AKMA 锚点功能（AAnF）实现，该功能可能运行在 5G 核心网边界之外。AAnF 被攻陷可能使恶意行为者利用漏洞，监控用户登录活动或未经授权地访问敏感通信内容。此外，订阅永久标识符（SUPI）对外部应用功能的暴露带来重大隐私风险，因为 SUPI 可被用于将用户的真实身份与其在线活动相关联，从而损害用户隐私。</span></p><p data-layout-id="422" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">为缓解这些漏洞，我们提出了一种名为 E2E-AKMA 的新协议，即使在 AAnF 被攻陷的情况下，也能在用户设备（UE）和应用功能（AF）之间建立具有端到端安全性的会话密钥。此外，该协议确保除 5G 核心网外，没有任何实体能将账户活动与用户的真实身份相关联。该架构保留了现有 AKMA 方案的优势，如无需复杂的动态秘密数据管理，且不依赖专用硬件（标准 SIM 卡除外）。实验评估表明，E2E-AKMA 框架相比原始 5G AKMA 方案的开销约为 9.4%，表明其具有部署的效率和实用性。</span></p><p data-layout-id="423" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-yueming.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_li-yueming.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="425"><span leaf="">164. Adversarial Patch EXterminator: Zero-Shot and Patch-Agnostic Defense Framework Against Adversarial Patch Attacks</span></h1><p data-layout-id="426" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Jiayimei Wang (香港城市大学); Tao Ni (阿卜杜拉国王科技大学); Guowen Xu (电子科技大学); Qingchuan Zhao and Cong Wang (香港城市大学)</span></p><p data-layout-id="427" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="428" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">对抗补丁攻击对现代计算机视觉系统构成严重威胁。尽管现有的防御方案试图通过开发可认证模型或补丁识别流程来缓解此类攻击，但它们普遍依赖先验知识或大量训练数据，在不同物理条件下鲁棒性不足，且在应对挑战性场景（例如微小、不规则或与背景高度融合的补丁）时性能有限。为解决这些局限性，我们提出了 APEX，一个零样本、补丁无关的三阶段对抗补丁防御框架。具体而言，APEX 首先通过边界框提取聚焦补丁区域，随后将基于互信息的模糊热力图与边缘感知的边界热力图相结合以定位对抗区域，最后利用结构引导的图像修复来恢复图像。我们在多个数据集和现有最先进防御方法上的实验表明，APEX 能有效防御各类对抗补丁（例如非自然主义、自然主义和红外图像补丁）。此外，APEX 在补丁定位方面展现出卓越能力，在不同环境（例如光照条件）和极端场景下保持高鲁棒性，并在物理世界场景中保护各类模型时表现出色。</span></p><p data-layout-id="429" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-jiayimei.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_wang-jiayimei.pdf</a></span></p><hr style="border-style: solid;border-width: 1px 0 0;border-color: rgba(0,0,0,0.1);-webkit-transform-origin: 0 0;-webkit-transform: scale(1, 0.5);transform-origin: 0 0;transform: scale(1, 0.5);"/><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="431"><span leaf="">165. Quantifying Large Language Model Attacks Through the Lens of Model Cognition</span></h1><p data-layout-id="432" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">作者：</span>Xiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai, Feiyue Xu, Sheng Hang, and Hanqing Hu (上海交通大学); Bin Benjamin Zhu (微软公司); Hongsheng Hu, Shi-Feng Sun, Dawu Gu, and Shuo Wang (上海交通大学)</span></p><p data-layout-id="433" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">摘要：</span></span></p><p data-layout-id="434" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">大语言模型（LLM）容易受到诱导有害内容的恶意输入的攻击。现有的安全机制（如关键词过滤或输出审核）在很大程度上忽略了模型内部动态。我们表明，与有害提示相关的安全特征在生成前的中间隐藏状态中可通过轻量级探针实现强可分性（准确率高达 99%），这表明即使模型输出合规内容，这些特征仍持久存在于内部。基于这一观察，我们引入了分层毒性探针和一种多层互补检测框架，融合来自不同深度的信号。我们的轻量级 Sentinel（&lt;5M 参数）相比生成级拒绝将假阴性率降低了一半，在对抗攻击下保持 94% 以上的检测准确率——而基线方法下降了 32%。Sentinel 在七个开源权重 LLM（1.5B→72B）和多个基准（I2P、SneakyPrompt、MMA、Labelled、PIJ、ChatAlpaca 和 Multi-turn Jailbreak）上的异构有害提示方面也优于 Llama-Guard-3-8B。除检测外，我们的方法提供了首张关于安全相关信号如何在 LLM 内部涌现、传播和衰减的定量分层图谱，实现了可解释的、由内而外的对齐和诊断。本文包含可能敏感和冒犯性的内容，包括但不限于 NSFW 材料、仇恨言论、歧视和其他有害文本。请读者审慎阅读。</span></p><p data-layout-id="435" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">PDF：</span><a href="https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-xiuming.pdf" target="_blank">https://www.usenix.org/system/files/conference/usenixsecurity26/sec26_prepub_liu-xiuming.pdf</a></span></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=dd471e21&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486124%26idx%3D2%26sn%3D725a2ac1a1c3c5eb5898347a88ea60bc">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 16 Jul 2026 20:22:00 +0800</pubDate>
    </item>
    <item>
      <title>面向自我演进的 Harness 工程研究综述</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486117&amp;idx=1&amp;sn=619c8f86e5c89b279f9de19a8dd72910</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>漏洞战争</span> <span>2026-07-11 09:59</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=b342df0b&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2syysYAaWpGw2zjkmxO1X5iaom1upXmWqT94FC0vuVlMLw4axTMicmgUCdzunAL4zZp8quoqNKgzxTYX7hotibon16A42bQ4K4a4Qc%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="0"><span leaf="">摘要</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="2"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">递归自我改进（Recursive Self-Improvement, RSI）是人工智能领域的长期愿景，其核心在于系统能够利用自身智能改进产生该智能的认知机制。近年来的实践表明，模型与真实世界之间的部署层——即 Harness——在 RSI 路径中扮演着与模型核心智能同等重要的角色。本文系统梳理了 Harness 工程的设计模式、优化方法及与模型权重的联合优化策略，涵盖工作流自动化、文件系统持久记忆、子agent并行执行等设计范式，以及上下文工程、进化搜索、自我改进循环等优化技术。在此基础上，本文进一步整合了自演进agent领域的理论框架（递归自指、轻量化范式、多agent协同进化）、标准化协议与基础设施、评测基准及错误进化风险等最新研究进展，并讨论了评估器脆弱性、上下文生命周期管理、多样性坍缩、奖励黑客等开放性挑战，最后对 Harness 层与核心智能的边界演化趋势进行了分析。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="3"><span leaf=""><span textstyle="" style="font-weight: bold;">关键词</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">：递归自我演进；Harness 工程；智能体系统；上下文工程；进化搜索；自演进Agent</span></span></p><p class="mp_profile_iframe_wrp" nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="0" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="5"><span leaf="">1 引言</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="6"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">递归自我改进的概念可追溯至 Good (1965) 提出的&#34;超智能机器&#34;设想——一种在所有智力活动中超越人类、并能设计更优机器以改进自身的系统 [1]。Yudkowsky (2008) 进一步将其明确为一个反馈回路：AI 利用当前智能改进产生该智能的认知机制本身 [2]。在现代 AI 语境下，这一反馈回路可以表现为模型直接重写自身权重，也可以更广泛地理解为模型改进训练流程和部署系统，从而催生性能更强的后继模型 3。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="7"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">值得关注的是，原始模型与真实世界上下文之间的部署层——即</span><span textstyle="" style="font-weight: bold;">Harness</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">——的重要性正日益凸显。Harness 是围绕基座模型构建的系统，负责编排执行流程，决定模型如何思考与规划、调用工具与执行动作、感知与管理上下文、存储制品以及评估结果。Claude Code 和 Codex 等编码agent产品的成功印证了这一点：模型的原始智能（即预训练后的评估表现）与 Harness 的工程质量共同决定了系统的实际能力上限。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="8"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">本文聚焦 Harness 工程的研究进展及其对 RSI 的贡献。与早期智能体框架&#34;agent = LLM + 记忆 + 工具 + 规划 + 动作&#34;的公式不同，Harness 工程在此基础上额外涵盖了工作流设计（如循环工程）、评估机制、权限控制和持久状态管理。它已不再仅仅是提示模板的组合，而更接近运行时和软件系统设计。</span></span></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="10"><span leaf="">2 Harness 设计模式</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="11"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 的设计应当刻意保持简洁和通用，以支持泛化能力。一种有效的策略是参照已有软件工程实践，从而受益于模型预训练阶段习得的知识。Harness 与操作系统之间存在强烈的类比关系：正如操作系统封装复杂逻辑的同时保持接口简洁，Harness 也应遵循类似原则。与此同时，配置、工具接口及其他协议有望逐步实现行业标准化。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="12"><span leaf="">2.1 工作流自动化</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="13"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">定义一个模型可以在其中操作、测试和迭代的工作流是实现自动化的关键设计。Karpathy 的 autoresearch 项目是该模式的典型实例 [5]。常见的工作流遵循一个目标导向循环：规划 → 执行 → 观察/测试 → 改进 → 再次执行，直至目标达成。在此过程中，模型可以主动向用户请求澄清任务规格或执行偏好。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="14"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="Codex agent循环简化示意" class="rich_pages wxw-img" data-ratio="0.4507042253521127" data-type="png" data-w="497" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002445" src="https://wechat2rss.xlab.app/img-proxy/?k=543cc66c&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sy84omlt8wvYmhSnlpghDt9A60g4FMj1lNGUfkEQnc08sKAIo1yBnR1iaQlCtIMILTMRXV5qPEpjSxmXO6kH02jTVasUMsn8I3I%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="15"><span leaf="">图 1. Codex agent循环简化示意：agent调用工具，工具响应影响模型的下一次生成。（图片来源：OpenAI Codex agent文章）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="16"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">工作流图的核心特征在于，模型通过&#34;agent运行时&#34;而非静态提示模板来分析自身轨迹和失败案例，并在此基础上迭代改进。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="17"><span leaf="">2.2 文件系统作为持久记忆</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="18"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">在长时程agent系统中，一个反复出现的模式是对丰富状态和制品的简洁控制。Harness 不应在上下文中承载整个工作流和所有日志，而应将持久状态保存在文件中。在长时程agent展开过程中，实验日志、代码差异、论文摘要、错误追踪和过往展开轨迹等制品的长度往往远超模型的上下文窗口。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="19"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">通过文件系统（通常借助</span></span><span leaf="" style="color:rgba(0, 0, 0, 0.9);font-size:17px;font-family:&#34;mp-quote&#34;, &#34;PingFang SC&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;line-height:1.6;letter-spacing:0.034em;font-style:normal;font-weight:normal;">bash</span><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">命令）进行读写和编辑是 LLM 的基础能力，因此以文件形式管理持久记忆能够自然地受益于核心模型能力的提升。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="20"><span leaf="">2.3 子agent与后端任务</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="21"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 可以生成多个子agent并行执行并监控后端任务。当主agent需要搜索多个假设、并发运行实验或委派隔离子任务时，这一模式尤为有用——它可以避免污染主上下文。父agent随后充当一个轻量进程管理器：启动任务、检查日志、取消失败运行、将结果合并回主agent线程。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="22"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">关键的设计选择在于使并行性显式化且可检查。如果子agent输出仅存在于瞬态聊天上下文中，它们很快会变得过时和隐蔽。若以文件、日志和状态记录的形式存储，模型则可以在中断后恢复，并对自身的执行历史进行推理。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="23"><span leaf="">2.4 编码agent Harness 案例分析</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="24"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">主流编码agent（Claude Code、Codex、OpenCode、Cursor 等）的核心接口已趋于稳定。它们通常采用如下循环结构：</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="25"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="编码agent Harness 循环" class="rich_pages wxw-img" data-ratio="0.33240740740740743" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002448" src="https://wechat2rss.xlab.app/img-proxy/?k=937d9d73&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxobbn5qKO6ypRosR9mhDNvsQ0qh6ucAo0xPHqSHANibUEUjAYarg6WUV1SfJ2BNXOWDrHK8wm6dcymANXKnEiaby35atE8LgCvw%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="26"><span leaf="">图 2. 编码agent Harness 循环结构</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="27"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">借助一组工具集，编码agent能够在给定代码库中开发和调试问题，类似于人类开发者使用 IDE 的方式。下表列出了编码agent常见的工具分组：</span></span></p><table style="box-sizing: border-box;border-collapse: collapse;border-spacing: 0px;width: 964px;overflow: auto;break-inside: auto;text-align: left;cursor: text;margin: 0px;padding: 0px;word-break: initial;white-space: pre-wrap;color: rgb(51, 51, 51);font-family: &#34;Open Sans&#34;, &#34;Clear Sans&#34;, &#34;Helvetica Neue&#34;, Helvetica, Arial, &#34;Segoe UI Emoji&#34;, &#34;SF Pro&#34;, sans-serif;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: normal;orphans: 2;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;"><thead><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;"><th style="box-sizing: border-box;padding: 6px 13px;font-weight: bold;border-width: 1px 1px 0px;border-top-style: solid;border-right-style: solid;border-left-style: solid;border-top-color: rgb(223, 226, 229);border-right-color: rgb(223, 226, 229);border-left-color: rgb(223, 226, 229);border-image: initial;border-bottom-style: initial;border-bottom-color: initial;margin: 0px;"><span cid="n31" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">工具组</span></span></span></th><th style="box-sizing: border-box;padding: 6px 13px;font-weight: bold;border-width: 1px 1px 0px;border-top-style: solid;border-right-style: solid;border-left-style: solid;border-top-color: rgb(223, 226, 229);border-right-color: rgb(223, 226, 229);border-left-color: rgb(223, 226, 229);border-image: initial;border-bottom-style: initial;border-bottom-color: initial;margin: 0px;"><span cid="n32" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">工具定义</span></span></span></th></tr></thead><tbody><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n34" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">文件系统</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n35" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">文件发现：</span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">glob</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">grep</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">ls</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">；文件读取：</span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">read</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">read_many</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">；文件修改：</span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">write</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">edit</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">multi_edit</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">apply_patch</span></code></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;background-color: rgb(248, 248, 248);"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n37" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Shell 执行</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n38" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">运行命令：</span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">bash</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">PowerShell</span></code></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n40" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">IO</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n41" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">lsp</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">；Git 工具：</span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">git_status</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">git_diff</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">git_commit</span></code></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;background-color: rgb(248, 248, 248);"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n43" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">外部上下文</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n44" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">MCP 工具, Skills</span></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n46" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">网络搜索</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n47" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">web_search</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">web_fetch</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, 浏览器工具</span></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;background-color: rgb(248, 248, 248);"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n49" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">制品生成</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n50" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">读取文档、图片；生成 HTML、图片</span></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n52" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">后端进程</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n53" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">CronCreate</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">CronDelete</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">CronList</span></code></span></span></td></tr><tr style="box-sizing: border-box;break-inside: avoid;break-after: auto;border: 1px solid rgb(223, 226, 229);margin: 0px;padding: 0px;background-color: rgb(248, 248, 248);"><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n55" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 79.925px;min-height: 10px;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">agent委派</span></span></span></td><td style="box-sizing: border-box;padding: 6px 13px;border: 1px solid rgb(223, 226, 229);margin: 0px;min-width: 32px;"><span cid="n56" mdtype="table_cell" style="box-sizing: border-box;display: inline-block;min-width: 1ch;width: 829.675px;min-height: 10px;"><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">spawn_agent</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">resume_agent</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">wait_agent</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">list_agents</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">close_agent</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">, </span></span><span md-inline="code" spellcheck="false" style="box-sizing: border-box;"><code style="box-sizing: border-box;font-family: var(--monospace);text-align: left;vertical-align: initial;border: 1px solid rgb(231, 234, 237);background-color: rgb(243, 244, 244);border-radius: 3px;padding: 0px 2px;font-size: 0.9em;"><span leaf="">interrupt_agent</span></code></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf=""> 等</span></span></span></td></tr></tbody></table><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="29"><span leaf="">2.5 Harness 层与核心智能的关系</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="30"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">预测 RSI 的未来在多大程度上依赖 Harness 工程是困难的，但近期路径不太可能从模型直接重写自身权重开始。本文认为，Harness 层与核心智能之间存在一个阶段性的协同演化关系，可从三个维度理解。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="31"><span leaf="">2.5.1 元方法论转向</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="32"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 工程将朝元方法论（meta-methodology）方向演化，即改进获取更优答案的机制本身，而非仅优化答案本身。在传统的 Harness 工程实践中，优化对象是具体任务的输出质量——通过调整提示模板、添加工具描述、修改工作流分支来提升模型在特定场景下的表现。元方法论转向意味着优化对象上移一个抽象层级：</span><span textstyle="" style="font-weight: bold;">Harness 系统自身成为优化目标</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="33"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一转向在前述各项研究中已具雏形。ACE 将上下文管理从手工规则转变为自动化的剧本演化；MCE 进一步将&#34;如何管理上下文&#34;的机制本身作为搜索对象；Meta-Harness 直接以编码agent为优化器，搜索决定信息存储、检索和呈现方式的代码；ADAS 和 AFlow 将工作流设计形式化为可通过元agent或 MCTS 自动搜索的优化问题。这些方法的共同特征是：</span><span textstyle="" style="font-weight: bold;">启发式的、人肉设计的规则逐步被通用的、自动化的机制取代</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。其本质是从&#34;工程师设计 Harness&#34;迁移至&#34;模型在设计空间中搜索 Harness&#34;，设计空间本身保持不变，但搜索主体和搜索效率发生了质变。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="34"><span leaf="">2.5.2 Harness 与模型智能的双向制衡</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="35"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">成熟的 Harness 系统与更强的模型智能之间存在双向制衡关系。一方面，成熟的 Harness 为模型自我改进循环中的自动研究提供了基础设施——AI Scientist 等系统已展示出从研究想法生成到论文撰写的全自动流水线能力 [10]。Harness 工程的质量直接决定了自改进循环能否闭合、能否在长时程任务中维持稳定。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="36"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">另一方面，更智能的模型能够防止 Harness 系统的过度工程化。当模型的核心推理能力足够强时，许多原本需要复杂 Harness 逻辑才能完成的任务（如上下文压缩、失败模式分析、工作流分支选择）可以由模型自身承担，从而简化 Harness 设计。这种制衡关系使系统保持可持续性——避免陷入&#34;Harness 越来越复杂、维护成本越来越高&#34;的正反馈失控。STOP 的实验结果为这一论点提供了经验佐证：递归自改进在 GPT-4 上有效，但在 GPT-3.5 和 Mixtral 等较弱模型上出现退化 [15]，表明基座模型的智能水平是 Harness 自改进机制发挥作用的前提条件。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="37"><span leaf="">2.5.3 内化与接口保留</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="38"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">最终，许多 Harness 改进可能会被内化为核心模型行为。当前由 ACE 策展器执行的结构化上下文更新、由 Self-Harness 实施的弱点修复、由进化搜索发现的工作流模式，在未来可能部分成为模型的原生能力——正如链式推理（Chain-of-Thought）曾是需要精心设计的提示技巧，而今已通过推理导向的强化学习内化为模型行为 [14]。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="39"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">然而，与外部上下文和工具的接口应当保留。模型无论多么智能，仍需接收关于目标、约束条件、上下文和评估方式的明确 specification。这一模式与提示工程的演化路径高度相似：随着指令调优和模型推理能力的提升，手动提示技巧的中心性显著降低，但指定目标、约束、上下文和评估的需求并未消失——它从&#34;教模型如何思考&#34;退化为&#34;告诉模型思考什么&#34; [6]。Harness 的演化可能遵循相同轨迹：</span><span textstyle="" style="font-weight: bold;">&#34;如何做&#34;的程序性技巧会被内化进模型权重，但&#34;做什么&#34;的接口 specification 将持续存在</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。这意味着 Harness 工程不会消亡，而是其重心从运行时逻辑设计转向接口与评估设计。</span></span></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="41"><span leaf="">3 Harness 优化</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="42"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 系统中优化对象的演进路径大致为：指令提示 → 结构化上下文 → 工作流 → Harness 代码 → 优化器代码。随着模型变得更智能、更强大，优化目标趋向更复杂的对象和更通用的方法。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="43"><span leaf="">3.1 上下文工程</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="44"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">简单地将所有工具响应和模型生成追加到上下文中，随着agent任务时程的增加会迅速失控。上下文管理是构建更结构化、更简洁上下文并管理持久状态的层。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="45"><span leaf=""><span textstyle="" style="font-weight: bold;">Agentic Context Engineering（ACE）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[7] 将上下文视为一份不断演化的&#34;剧本&#34;（playbook），而非不断增长的提示。ACE 包含三个组件：</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="46"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">生成器：参考要点生成任务轨迹。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="47"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">反思器：从成功和失败轨迹中提炼洞见。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="48"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">策展器：以增量、条目化的方式更新结构化上下文。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="49"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="ACE 框架" class="rich_pages wxw-img" data-ratio="0.3675925925925926" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002447" src="https://wechat2rss.xlab.app/img-proxy/?k=95cd7584&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2swMwhyY4stWSudaJ0yPNfO7eZqtofL3tSyFQbsnicsF2oWjicGJ4MQNQGmibzs2Rr7ZgIFdHic6gHeU4r2CFNeqss1CCwNAX0MJia48%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="50"><span leaf="">图 3. Agentic Context Engineering (ACE) 框架。（图片来源：Zhang et al. 2025）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="51"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">为防止迭代重写中的上下文坍缩和简洁性偏差，ACE 的策展器不重写完整的提示块，而是输出结构化条目（标识符, 描述），通过确定性逻辑合并到结构化上下文日志中，并周期性地进行精炼和去重。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="52"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">ACE 从展开中学习洞见，推动自管理记忆的发展，但其更新规则和工作流仍是手工设计的。</span><span textstyle="" style="font-weight: bold;">Meta Context Engineering（MCE）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[8] 将机制（如何管理上下文）与制品内容（上下文中有什么）分离，在元优化层面运行技能演化，在基础层面运行上下文优化。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="53"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">MCE 中的技能 s∈Ss \in S 定义了上下文函数 cs=(ρs,Fs)c_s = (\rho_s, F_s)，将输入 xx 映射到上下文 c=Fs(x;ρs)c = F_s(x; \rho_s)，其中 ρs\rho_s 为静态组件（提示、知识库、代码库），FsF_s 为动态算子（搜索、选择、过滤、格式化）。双层优化公式为：</span></span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="54"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.4069400630914827" data-s="300,640" data-type="png" data-w="317" type="block" data-imgfileid="100002466" src="https://wechat2rss.xlab.app/img-proxy/?k=114dd24e&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2swS9bCTxtv2HOLw8vU2XeSVuTgJUAjCmsQWQdnyFH5cGWCZEAgd7HsXaCLg4kvrkcGGH2icL8TkjQ7Gl53aaG9dHF07qbRiah7EQ%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="55"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="MCE 框架" class="rich_pages wxw-img" data-ratio="0.5287037037037037" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;" data-imgfileid="100002450" src="https://wechat2rss.xlab.app/img-proxy/?k=100d601f&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sygDYt8C4vXR4O6ArjSE9RVyUHF8ExYUSwbdQTbnib2bYheX15zZVNVxIZ5m6Y6Ria8I9ODAYEtOJxnSicDBNS0zic1tZ4LS6VichAA%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="56"><span leaf="">图 4. Meta Context Engineering (MCE) 框架：元级技能演化搜索上下文管理机制，基础级优化任务上下文。（图片来源：Ye et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="57"><span leaf=""><span textstyle="" style="font-weight: bold;">Meta-Harness</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[9] 将优化目标进一步深入：被优化的对象是决定什么信息应被存储、检索和呈现给模型的代码本身。其命名中的&#34;Meta&#34;意味着它是优化 Harness 的 Harness。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="58"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="Meta-Harness 外循环优化算法" class="rich_pages wxw-img" data-ratio="0.4388888888888889" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002446" src="https://wechat2rss.xlab.app/img-proxy/?k=f0d867a4&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2sxKCk3GE1ibia1HbTv3RX9hdKjno1gNosne4ztkC9ayNo5TDAt8VTLqKw4MRCibX6lZBNADvQem1Dg1eo2ibJfK6qHIYib4HiaDcr1SI%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="59"><span leaf="">图 5. Meta-Harness 外循环优化算法。（图片来源：Lee et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="60"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">提出新 Harness 的提议者本身是一个编码agent，最终输出是帕累托前沿上的一组 Harness 候选。整个执行历史通过文件系统访问，编码agent使用</span><span textstyle="" style="background-color: rgba(0, 0, 0, 0.9);color: rgb(43, 119, 191);">grep</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">或</span><span textstyle="" style="background-color: rgba(0, 0, 0, 0.9);color: rgb(43, 119, 191);">cat</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">等命令阅读历史，而非将所有内容塞入单个提示上下文。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="61"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="Meta-Harness 性能" class="rich_pages wxw-img" data-ratio="0.3212962962962963" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002451" src="https://wechat2rss.xlab.app/img-proxy/?k=2fa239c5&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2swRUSuJYEyan5P0YiarIP2o1a35rqBuP4ul9iafm4P7d7NNrkJfeosmrtiaYdZXxb4C528ygE2UU80flialZJmwCxTpAJnBqKDibbWw%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="62"><span leaf="">图 6. Meta-Harness 在文本分类（左）和 TerminalBench-2（右）上的性能表现。TerminalBench-2 实验从 Terminus-KIRA 和 Terminus-2 两个强 Harness 初始化搜索。（图片来源：Lee et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="63"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">核心启示在于：一旦 Harness 设计成为可执行的搜索空间，强大的编码agent就能利用人类工程师所使用的同一设计空间。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="64"><span leaf="">3.2 工作流设计</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="65"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">工作流设计可由领域专家手工构建。以自动研究为例，</span><span textstyle="" style="font-weight: bold;">AI Scientist</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[10] 构建了一条从提出研究想法、编写代码、运行实验、分析结果、撰写手稿到同行评审的完整流水线。</span><span textstyle="" style="font-weight: bold;">ScientistOne</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[11] 将可验证性作为核心设计约束，每条声明（引用、数值、方法论、结论）都必须追溯到证据源，并通过 Chain-of-Evidence 审计。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="66"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="AI Scientist 流水线" class="rich_pages wxw-img" data-ratio="0.7751479289940828" data-type="png" data-w="1014" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002452" src="https://wechat2rss.xlab.app/img-proxy/?k=4855bc9a&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxVqiaSFiaanpDfLk3YLKD2mXmC25PwEKf7TicYLxHC3wzSVicibQQRBTouFrFhL08iajFFmOsdlzxlXsDCUZDRBXnWnfkMdf15YgcyE%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="67"><span leaf="">图 7. AI Scientist 流水线：从想法生成、实验、论文撰写到评审。（图片来源：Lu et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="68"><span leaf=""><span textstyle="" style="font-weight: bold;">Autodata</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[12] 作为数据科学家agent，旨在生成训练和评估数据。主agent管理挑战者（提出问题）、弱求解器、强求解器和验证器/裁判，目标是合成&#34;恰好合适&#34;难度的数据——即强求解器成功而弱求解器失败的任务。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="69"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="Autodata 工作流" class="rich_pages wxw-img" data-ratio="0.5583333333333333" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002453" src="https://wechat2rss.xlab.app/img-proxy/?k=e93c0bc3&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2syBhK3gs8Zvu8FJoVzaLo9dTWa9JPYDXBuKQuJvCgIOZyndOPPsbwsicfzTia6ttkE2z69eXdiajnPdqgt3OPf2fT2icQN006OYYGo%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="70"><span leaf="">图 8. Autodata agent工作流设计：围绕挑战者、求解器和验证者角色生成合成数据。（图片来源：Kulikov et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="71"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">工作流设计空间极其庞大，可将其视为搜索问题。</span><span textstyle="" style="font-weight: bold;">Automated Design of Agentic Systems（ADAS）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[13] 将agent设计本身形式化为优化问题，由元agent提出新的工作流设计。初始化一个包含简单agent（如 CoT、self-refine）的档案库后，元agent受档案库中已有方案启发，以代码形式编程新agent，经自精炼检查新颖性后评估，成功者加入档案库。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="72"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="ADAS 示意" class="rich_pages wxw-img" data-ratio="0.5703703703703704" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002454" src="https://wechat2rss.xlab.app/img-proxy/?k=0c071401&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxyhQM6omm9jRWFGRrovhajNAzn42SVVgC4eA9m1l3FZlXAIianibzAn0DKkYsOOnKZ2BHzwPmmLR5j7cribCEtbXSEedSVt1iayww%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="73"><span leaf="">图 9. Automated Design of Agentic Systems (ADAS) 示意图。（图片来源：Hu et al. 2025）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="74"><span leaf=""><span textstyle="" style="font-weight: bold;">AFlow</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[14] 将agent工作流表示为图，节点为 LLM 调用动作，边为代码实现的逻辑操作。工作流优化依赖蒙特卡洛树搜索（MCTS），在 QA、代码和数学任务上表现出对手工工作流和 ADAS 的改进。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="75"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="AFlow 优化过程" class="rich_pages wxw-img" data-ratio="0.6240740740740741" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002455" src="https://wechat2rss.xlab.app/img-proxy/?k=0ba06323&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2syWmbXDdR8DpPTm81NcFYicd7HB2zavvfyv05IsrxjUgKk041ribBExic9UM7tkexv5j1icZlJnHQUHJTAlgiag8T34icuzZNiaQlQNF4%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="76"><span leaf="">图 10. AFlow 在工作流候选树上的优化过程。（图片来源：Zhang et al. 2025）</span></figcaption><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="77"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="AFlow 实验结果" class="rich_pages wxw-img" data-ratio="0.3296296296296296" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002460" src="https://wechat2rss.xlab.app/img-proxy/?k=c81609e4&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2syQ7DdeIiclADU3FURSm3WSEQ5iccKjtgPegRV1DLNIiaGEgEMBYicibS2uDJCWLyeEYjicuKePrgFn4xfl7VFiabS28P9ZkRA3qrRxhs%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="78"><span leaf="">图 11. AFlow 与手工方法和 ADAS 的实验对比。（图片来源：Zhang et al. 2025）</span></figcaption><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="79"><span leaf="">3.3 自我改进的 Harness</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="80"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">上下文工程和工作流设计仅是 Harness 的局部。完整的 Harness 搜索需要同时优化上下文管理逻辑、工作流、权限控制等多个组件。代码是定义程序和系统的通用语言——一个 Harness 本质上是一段协调提示、工具调用、子agent、控制流、记忆和工作流逻辑的代码。如果 LLM 能够优化执行agent的代码，它就能访问比手写提示大得多的设计空间。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="81"><span leaf=""><span textstyle="" style="font-weight: bold;">Self-Taught Optimizer（STOP）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[15] 是递归支架改进的早期实例。种子改进器 I_0 在步骤 t=0接收初始解 s、效用函数 u 和黑盒语言模型 M，返回改进解 s&#39;。STOP 的目标不是直接改进 s，而是改进改进器 I 本身。元效用定义为改进器函数 I 在下游任务集合 D 上的平均效用：</span></span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="82"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.18128654970760233" data-s="300,640" data-type="png" data-w="342" style="margin-top: 17px;margin-bottom: 17px;float: none;" type="block" data-imgfileid="100002463" src="https://wechat2rss.xlab.app/img-proxy/?k=07ba1fc0&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2szHtSicdJI5hiclTpA1AEx0UZvicPDx6icQNrKS2HUDdgy0AuM54kMsqHUv75ibBc26ic3GHDpkNPcgJRlrmicBzb0QULy39WUCMjekRY%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="83"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">通过递归自改进更新：</span></span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="84"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.15979381443298968" data-s="300,640" data-type="png" data-w="194" type="block" data-imgfileid="100002464" src="https://wechat2rss.xlab.app/img-proxy/?k=aa75edf6&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxerckhN2IKXDVhGzoBbkV8ibwf8aORW2kmgphDjg6ofmPNSUIia8ziaicKWpuicaicdh3ibPDD63ibUjFvqBhQNPLKDtVuxKKSSFzjH9g%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="85"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="STOP 算法" class="rich_pages wxw-img" data-ratio="0.37962962962962965" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;" data-imgfileid="100002458" src="https://wechat2rss.xlab.app/img-proxy/?k=fb8dd1f8&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sx40FzpcJVuAeu0HA08MsWZ4dZz7Vlq2g2K0JmKVYLiaMDicMCLYJb3tPM4ia7vmW2khbJZUzZQHdzZhbTFYgTg7ZPPic94RAzIicls%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="86"><span leaf="">图 12. Self-Taught Optimizer (STOP) 算法。（图片来源：Zelikman et al. 2023）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="87"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">实验中，STOP 在 GPT-4 上实现了跨迭代平均下游性能的提升，但在 GPT-3.5 和 Mixtral 等较弱模型上出现退化。这表明递归结构本身不够——基座模型必须具备足够的智能才能改进机制。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="88"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="STOP 发现的自改进策略" class="rich_pages wxw-img" data-ratio="0.22592592592592592" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002457" src="https://wechat2rss.xlab.app/img-proxy/?k=d1748ab3&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxpbOnMuZDKO0h8ZavMbwnGAEOYBPhA3aA0GJwdfDEJ3D4kDrmMLK5aG5DDFT5TzSo71SvR4oyQibeoSrtGHQvkohhVpF6VFryI%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="89"><span leaf="">图 13. STOP 发现的自改进策略示例。（图片来源：Zelikman et al. 2023）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="90"><span leaf=""><span textstyle="" style="font-weight: bold;">Self-Harness</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[16] 依赖 LLM agent通过&#34;提出-评估-接受&#34;循环改进自身 Harness，包含三个阶段：</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="91"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">弱点挖掘：将失败聚类为验证器基础的失败模式，需记录包含终端验证器级原因、相关agent行为的因果状态以及轨迹暴露的抽象agent机制的丰富信息。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="92"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 提议：基于挖掘的失败模式提出有界 Harness 编辑，需考虑当前 Harness 的可编辑表面、验证器基础的失败模式、应保留的通过行为记录以及先前尝试的编辑摘要。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="93"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">提议验证：在保留集上评估候选编辑，仅在无回归时接受并合并。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="94"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="Self-Harness" class="rich_pages wxw-img" data-ratio="0.7203703703703703" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002459" src="https://wechat2rss.xlab.app/img-proxy/?k=fbc13c61&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2szfSRVJXQftcoUWsWz4dmyAfibcSAGibk0OS89VC8uDiaWJ77caKSLvOysY7icVNsErTUa1h5qFSnwawtrZpEMQvcEU2OlGOw798fI%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="95"><span leaf="">图 14. Self-Harness 通过弱点挖掘、有界 Harness 提议和验证循环来更新 Harness。（图片来源：Zhang et al. 2026）</span></figcaption><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="96"><span leaf="">3.4 进化搜索</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="97"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">进化搜索受自然选择启发，通过变异种群中的解并保留高适应度个体来优化。当搜索空间庞大或形状不规则、难以用梯度直接优化但易于评估解时，进化搜索尤为适用——Harness 搜索恰好符合这些条件。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="98"><span leaf=""><span textstyle="" style="font-weight: bold;">AlphaEvolve</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[17] 作为编码agent进化搜索系统，存储候选程序池并提示冻结的 LLM 生成改进差异。设计上的关键细节包括：提示包含父程序、结果和指令；代码改进区域以</span></span></p><h1 style="font-size: 20px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf="">EVOLVE-BLOCK-START和<span textstyle="" style="background-color: rgba(0, 0, 0, 0.9);color: rgb(43, 119, 191);"># EVOLVE-BLOCK-END</span>显式标记；元提示与指令和上下文共同进化。</span></h1><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="99"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="AlphaEvolve" class="rich_pages wxw-img" data-ratio="0.6222222222222222" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002456" src="https://wechat2rss.xlab.app/img-proxy/?k=f79a873e&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sxZyQs2kVGa96HXTGRo2gOtXqey7HtEMqIPbJl4WGrqGPCFrkvezyWT48CW2bNDKt5MOjuUYscpVhA4B8kWibzT2fZ9b0v0BaKY%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="100"><span leaf="">图 15. AlphaEvolve 工作原理。（图片来源：Novikov et al. 2025）</span></figcaption><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="101"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="AlphaEvolve 消融实验" class="rich_pages wxw-img" data-ratio="0.3527777777777778" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002462" src="https://wechat2rss.xlab.app/img-proxy/?k=80770db0&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2syeicXPffdezhaaibJUgCgDeyeQUhFbDwedyicG3FkZaXsItgC01n3AiawibibLjQq3EybCzvgE4hgNEpPuhNW86nR7mOakC0cu41j44%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="102"><span leaf="">图 16. AlphaEvolve 各设计组件的消融实验结果。（图片来源：Novikov et al. 2025）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="103"><span leaf=""><span textstyle="" style="font-weight: bold;">Darwin Gödel Machine（DGM）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[18] 明确以可编辑的 Harness 代码仓库为进化目标。从池中一个编码agent开始，每次迭代按性能正比、子代数量反比的概率选择父agent，由其检查自身基准评估日志并提出 Harness 代码库改进。以 Claude 3.5 Sonnet 为基座 LLM，DGM 发现的agent在 SWE-bench Verified 上从 20% 提升至 50%，在 Polyglot 上从 14.2% 提升至 30.7%。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="104"><span leaf="">3.5 与模型权重的联合优化</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="105"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 进化改变的是模型周围的非参数化系统。为实现完全的自我改进，模型也可以同时更新自身权重。</span><span textstyle="" style="font-weight: bold;">SIA</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[19] 是将 Harness 改进和模型参数更新纳入同一优化循环的早期尝试，包含三个组件：</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="106"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">元agent：提出初始 Harness。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="107"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">任务agent：执行任务。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="108"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">反馈agent：根据近期轨迹决定下一步是更新 Harness 还是模型权重。</span></span></p><div style="color: rgba(0, 0, 0, 0.9);text-align: start;font-size: 17px;font-weight: 400;line-height: 1.8;margin-bottom: 24px;" data-layout-id="109"><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" alt="SIA" class="rich_pages wxw-img" data-ratio="0.39444444444444443" data-type="png" data-w="1080" style="box-sizing: border-box;border-width: 0px 4px 0px 2px;border-top-style: initial;border-right-style: solid;border-bottom-style: initial;border-left-style: solid;border-top-color: initial;border-right-color: transparent;border-bottom-color: initial;border-left-color: transparent;border-image: initial;vertical-align: middle;max-width: 100%;image-orientation: from-image;cursor: default;display: block;height: auto;margin: auto;margin-top: 17px;margin-bottom: 17px;float: none;" data-imgfileid="100002461" src="https://wechat2rss.xlab.app/img-proxy/?k=b1c2269f&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sygBqYM9YwvFm3GdCsfo2mIxfftvUzMyRWyu7aR0vZPqQKMIs3pwreBVAFkqvJQyJfMlIAQib9aExWsJNJ3oX7qEccRiaqqQbFqI%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="110"><span leaf="">图 17. SIA 中反馈agent决定下一次迭代类型。（图片来源：Hebbar et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="111"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">SIA 的实验存在一些混淆因素（如任务agent使用的模型远弱于元agent和反馈agent的模型），其结果应被视为初步证据。训练稳定性和 Goodhart 效应等挑战仍然存在。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="112"><span leaf="">3.6 面向自演进的系统基础设施</span></h2><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="113"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.5211693548387096" data-s="300,640" data-type="png" data-w="992" style="margin-top: 17px;margin-bottom: 17px;float: none;" type="block" data-imgfileid="100002465" src="https://wechat2rss.xlab.app/img-proxy/?k=ccc0aff8&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2swI90RiaWR6UEXt4kK8gt47QjhkoH8ncLtEYB6FGwgKics9BgpKG1EO8bTJm3ibiaOMpXwhTj0ZqeqFVpzkz3Bs6XwPmBFXXqhBtyE%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><figcaption style="text-align: start;font-size: 14px;font-weight: 400;color: rgba(0,0,0,0.3);line-height: 1.8;margin-bottom: 24px;" data-layout-id="114"><span leaf="">图 18. AREAL2.0在线强化学习工作流示意图：现有智能体服务保留其原有的规划、工具执行、沙箱和记忆模块，同时将LLM API调用重定向至网关，随后路由器通过数据agent将请求路由至智能体计算工作节点，记录已部署的轨迹数据，并将推理服务与在线强化学习训练相连，以实现策略模型权重的更新。（图片来源：Yan et al. 2026）</span></figcaption><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="115"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">前面介绍的优化方法大多在实验室或单任务场景中验证，而企业级大规模agent服务面临的挑战有所不同。Yan et al. (2026) 在 AReaL 2.0 [28] 中指出，自演进agent在企业级部署中受阻的根本原因不在于 RL 算法本身，而在于缺乏支撑在线agent强化学习的系统基础设施。当前部署的 LLM agent在权重、系统提示、工具库和上下文 Harness 均在部署时冻结，任何改进都依赖人工数据收集、离线微调、范式修改和重新部署的手动循环。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="116"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">该论文提出下一代agent RL 系统须围绕三大支柱协同设计：</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="117"><span leaf=""><span textstyle="" style="font-weight: bold;">支柱一：agent轨迹数据协议（Agent Trajectory Data Protocol, ATDP）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。现有agent日志通常记录提示、补全、工具调用、延迟和错误等信息，适用于调试但不适用于在线 RL。学习就绪的轨迹必须保留步粒度的决策过程：agent可用的观察、相关内部或 Harness 状态、选择的动作、动作结果、延迟到达的奖励或批评信号，以及模型版本、工具模式、租户、成本和治理状态等元数据。ATDP 需厂商中立、框架无关，支持延迟反馈和奖励增强，保留版本化溯源以支持反事实重放，并从一开始就编码隐私、访问控制、保留期限和训练资格字段。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="118"><span leaf=""><span textstyle="" style="font-weight: bold;">支柱二：企业级agent数据agent（Agentic Data Proxy）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。ATDP 规定须表示什么，数据agent规定如何在异构企业环境中产生这些表示。由于部署的agent不会基于单一框架或提供商，数据agent须在稳定的执行边界处拦截：LLM 调用、工具调用、检索调用、记忆读写、文件或浏览器操作、审批事件、用户纠正和最终任务结果。其角色不仅是导出轨迹，而是通过脱敏敏感字段、执行访问控制和保留策略、附加溯源、采集弱信号和延迟奖励，将生产工作转化为受治理的学习材料。关键在于区分确定性重放、近似重放和不可重放事件——企业学习系统须能回答&#34;在不同提示、模型检查点、记忆条目或工具模式下，agent是否仍会成功&#34;。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="119"><span leaf=""><span textstyle="" style="font-weight: bold;">支柱三：统一agent演进控制平面（Agent Evolution Control Plane）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">。自演进不应等同于盲目更新模型权重。部署的agent是一个由基座 LLM、上下文 Harness、记忆、工具和护栏组成的复合策略，不同的失败需要不同的干预表面：反复缺失的事实可能需要记忆插入，工具路由失败可能需要 Harness 或模式编辑，可复用的程序性失败可能需要技能补丁，跨租户、任务和工具配置持续出现的广泛失败则可能需要通过 SFT、偏好优化、在线 RL 或蒸馏来更新模型权重。控制平面将自演进视为一个受治理的决策问题——给定一组 ATDP 轨迹、轨迹统计、评估器分数、用户纠正率、工具失败聚类、成本信号、安全约束和分布漂移指标，在记忆更新、技能更新、Harness 编辑、工具模式变更、策略更新、回滚或无操作之间做出选择。每项选定的干预须经过先重放评估、离线回归测试、租户感知安全检查和版本化回滚。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="120"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AReaL 2.0 作为该愿景的一个限定分支原型，聚焦于在线策略 LLM 权重更新这一代表性演进路径，展示了如何将现有离线后训练 RL 框架重组为面向agent服务的在线 RL 循环：部署的agent服务可将 LLM 推理调用重定向至 AReaL 2.0 管理的agent计算工作节点，交互轨迹则由 RL 训练管线捕获和消费 [29]。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="121"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一工作将 Lilian Weng 所述的优化对象演进路径（指令提示 → 结构化上下文 → 工作流 → Harness 代码 → 优化器代码）进一步延伸至</span><span textstyle="" style="font-weight: bold;">系统级的自动决策层</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">——不仅优化某一特定表面，而是自动决定该优化哪一个表面。SIA 的反馈agent是这一思路的雏形，而 AReaL 2.0 的控制平面将其系统化为企业级的多表面治理框架。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="122"><span leaf="">3.7 自演进agent的理论框架与生态全景</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="123"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">自演进agent作为一个独立研究方向，已形成从理论框架到关键实现技术再到评测基准的初步生态。两篇 2025–2026 年的综述为该领域建立了统一的术语体系和分类标准。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="124"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Gao et al. (2026) [30] 在首篇系统性综述中围绕&#34;进化什么&#34;&#34;何时进化&#34;&#34;如何进化&#34;三大维度完成全领域梳理，覆盖模型、记忆、工具、架构四类可进化组件，区分测试内迭代与长期在线演化等进化阶段，对比标量奖励、文本反思、单/多智能体协同等进化机制。Fang et al. (2025) [31] 则提出了自进化智能体的统一反馈循环抽象框架，包含输入、智能体、环境、优化器四大核心模块，并提出了自进化&#34;三定律&#34;：安全适应、性能保持、自主进化——为系统的安全可控提供了设计原则。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="125"><span leaf="">3.7.1 递归自指框架</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="126"><span leaf=""><span textstyle="" style="font-weight: bold;">Gödel Agent</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[32] 受哥德尔机理论启发，提出了首个工程可落地的全自指递归自进化框架。区别于固定架构、固定元学习算法的传统agent，Gödel Agent 可在运行时读取并动态改写自身的决策逻辑、工具调用策略与优化算法，无需预设参数搜索空间。在数学推理、科学问答和代码生成任务上，它超越了 ReAct 和 MetaAgent Search 等基线，能自主演化出启发式搜索和回溯等复杂求解策略。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="127"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">与本文前述的 DGM（Darwin Gödel Machine）[18] 相比，Gödel Agent 侧重单agent的递归自指修改，而 DGM 引入了种群层面的开放式进化搜索，二者从不同角度推动了递归自改进的理论边界。后续工作</span><span textstyle="" style="font-weight: bold;">Polaris</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[33] 将此框架推广至小语言模型（7B），通过经验抽象（experience abstraction）与最小补丁修复（minimal code patch repair），在多项基准测试上实现了超过大模型的性能，表明递归自进化并非大模型专属。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="128"><span leaf="">3.7.2 轻量化自进化范式</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="129"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">并非所有自进化路径都需要修改模型权重或重写代码。</span><span textstyle="" style="font-weight: bold;">Reflexion</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[34] 开创了无微调的轻量化自进化范式：agent对任务中的错误进行语言层面的自我反思，将反思文本保存在情景记忆缓冲区中，以指导后续尝试。这一方法计算成本极低，不改动模型权重即可实现迭代提升，已成为几乎所有自进化工作的对比基线。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="130"><span leaf=""><span textstyle="" style="font-weight: bold;">MetaAgent</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[35] 聚焦工具侧自进化，以&#34;做中学&#34;为原则，从最小化工作流开始，在任务执行中持续反思与验证，将经验浓缩为可迁移文本动态融入后续上下文，同时通过管理工具使用记录构建持久的内部知识库与工具体系，实现零模型微调的自进化。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="131"><span leaf=""><span textstyle="" style="font-weight: bold;">Agent0</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[36] 提出了零外部数据的双agent协同自进化范式：课程agent持续生成难度递增的前沿任务，执行agent借助工具完成求解，执行能力提升倒逼课程agent产出更复杂任务。基于 Qwen3-8B 的实验显示，该方案将数学推理性能提升 18%、通用推理性能提升 24%，彻底摆脱了对人工数据集的依赖。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="132"><span leaf="">3.7.3 多agent去中心化协同进化</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="133"><span leaf=""><span textstyle="" style="font-weight: bold;">MorphAgent</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[37] 针对中心化调度瓶颈，提出去中心化协同自进化框架：agent可自主演化专属角色档案、动态调整集群协作拓扑，无需中心管控即可在开放环境中自主分工、互相学习迭代。其分布式经验沉淀机制实现了跨agent能力迁移，适用于机器人集群、分布式 AI 协作平台等场景。这一工作填补了集群层面自主分工与动态角色迭代的技术空白。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="134"><span leaf="">3.7.4 标准化协议与基础设施</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="135"><span leaf=""><span textstyle="" style="font-weight: bold;">Autogenesis Protocol（AGP）</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[38] 从系统架构视角提出了双层标准化自进化协议：资源基底协议层（RSPL）统一建模提示词、agent、工具、记忆等全量可进化资源；自进化协议层（SEPL）定义带审计、回滚能力的闭环进化算子，实现进化对象与进化逻辑解耦。基于 AGP 构建的多agent系统在长时序、多工具复杂基准上优于基线，为工业化可管控的自进化agent提供了统一底层标准。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="136"><span leaf=""><span textstyle="" style="font-weight: bold;">AgentGym</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[39] 推出了领域首个统一多环境agent训练与自进化开源平台，集成 14 类异构交互环境、标准化专家轨迹数据集 AgentTraj 与统一评测基准 AgentEval，配套跨环境自进化算法 AgentEvol，解决了环境碎片化和缺少通用评测工具的行业痛点。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="137"><span leaf="">3.7.5 评测基准：从实验室到经济价值场景</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="138"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">自进化agent的评测本身是一个独立挑战。</span><span textstyle="" style="font-weight: bold;">GDPevo</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[40] 据称是首个在具有真实经济价值（GDP 相关）的任务上专门评估agent自进化能力的基准。它覆盖 CRM、ERP 和金融三大场景共 120 个真实企业任务，采用&#34;规则杂交&#34;技术——将复杂业务逻辑拆分为元规则分散藏入训练集，再将其重新组合为测试题——以区分&#34;背答案&#34;和&#34;学规则&#34;。GDPevo 使用确定性规则打分器而非 LLM-as-a-Judge，保证分数可复现且失败可追溯。实验显示，当前主流agent（Claude Code、Codex）的自进化能力可将测试准确率提升约 17–22%，同时降低 Token 消耗。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="139"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">该基准揭示了一个重要信号：当前agent已具备一定的自进化能力，但评估方法的成熟度仍是制约研究进展的瓶颈。GDPevo 采用端到端全自动的基准构建流程以对抗数据泄露，这一思路与 Loop Engineering 的理念一脉相承——当出题速度超过模型记忆泄露答案的速度时，基准始终保持有效。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="140"><span leaf="">3.7.6 错误进化风险</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="141"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">自进化并非总是朝正确方向演进。</span><span textstyle="" style="font-weight: bold;">Misevolution</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[41] 研究了自进化agent的&#34;错误进化&#34;风险：agent可能从经验中学到错误模式并将其固化为行为策略，导致性能退化而非提升。这一问题与本文第 4.5 节讨论的奖励黑客密切相关，但更为隐蔽——错误进化不一定源于奖励信号被优化，而可能源于agent从有限经验中归纳出错误的因果推断。这要求自进化系统具备错误检测与回滚机制，而非无条件接受所有&#34;学到的&#34;改进。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="142"><span leaf="">3.7.7 Skill 作为 Harness 自进化的初级形式与产业实践</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="143"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">崔添翼（DeepSeek Harness 负责人）在转发翁荔博文时提出了一个重要的分层观察：</span><span textstyle="" style="font-weight: bold;">Skill 是 Harness 自进化中比较初级的一种形式，即从 prompt 层面进行自进化</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[42]。这一观点为 Harness 优化路径提供了一个更精细的梯度划分——从 prompt 层的 Skill 自进化，到上下文工程的结构化演化，再到 Harness 代码层面的重写，自进化的深度逐级递增。Skill 层面的自进化门槛最低、见效最快，是 Harness 自进化最务实的入口。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="144"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一判断已获得产业界的验证。DeepSeek 在 2026 年中期已形成&#34;模型 + Harness&#34;双轮驱动的研发模式，由专职团队推进 Agent 架构、上下文管理、多智能体调度和工具集成方向的工程化 [42]。这一模式并非 DeepSeek 独有——头部 AI 公司均在向&#34;模型 + 模型外系统&#34;的双轮配置迁移：OpenAI 以 GPT 模型配合 Function Calling 与 GPT-Live 工具 Harness；Anthropic 以 Claude 模型配合 Claude Code 编程 Harness 与 MCP 工具协议；Google 以 Gemini 模型配合 Agent Kit 多模态 Harness。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="145"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">从产业边界视角审视，当前主流编码代理产品（Cursor、Claude Code、Devin、通义灵码等）本质上都是 Harness 公司——它们的核心竞争力不在基座模型，而在于 Harness 工程质量和工具生态的成熟度。Harness 正在成为&#34;模型变现的最后一公里&#34;：同一个基座模型放入不同 Harness 中可能表现出完全不同的能力，这一观察已从少数人的经验判断演变为产业共识 42。GitHub Trending 数据也佐证了这一趋势——Agent/Harness 类项目持续占据前列，而纯模型项目逐渐退出热门榜单。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="146"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">翁荔（前 OpenAI 安全副总裁、Thinking Machines Lab 联合创始人）与崔添翼（DeepSeek Harness 负责人）在 Harness 自进化方向上的共识具有标志性意义。两位横跨中美 AI 研究圈的代表性人物罕见达成一致，表明&#34;自进化先从 Harness 开始&#34;已不是个别观点，而是产业级判断 42。这与本文 2.5.1 节论述的元方法论转向形成呼应——Harness 工程的竞争重心正从&#34;训练侧&#34;（谁的模型更大）转向&#34;工程侧&#34;（谁的 Harness 更优），AI 产业的护城河正在从&#34;模型权重&#34;向&#34;Harness 工程质量 + 工具生态 + 上下文管理能力&#34;迁移。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="147"><span leaf="">3.7.8 自进化的使能条件与耦合动力学</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="148"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Apodex（陈天桥创立的 AI 公司，专注&#34;发现模型&#34; Discovery Model）的两位首席科学家杜少雷与 Beibin Li 在 2026 年 7 月的深度对话中 [44]，从工业实践角度揭示了自进化闭环的若干使能条件和耦合关系，为本文的理论框架提供了实践佐证。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="149"><span leaf=""><span textstyle="" style="font-weight: bold;">编码能力作为自进化的底座。</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">训练一个模型本质上是数据、基础设施和算法三件事，而这三者都高度依赖代码——合成数据生成、自然语言数据清洗、训练代码和底层 infra 代码均以代码形式存在。一旦模型的编码能力足够强（如 Anthropic 报告其 80% 的代码由模型自身编写），用它来进行自我提升就成为非常自然的事情 [44]。这一观察为本文 3.7.7 节的 Skill 分层观点提供了更底层的解释：Skill 层面的自进化之所以是最务实的入口，正是因为它直接利用了编码这一使能能力。换言之，</span><span textstyle="" style="font-weight: bold;">编码能力是自进化阶梯的地基</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">——没有足够强的编码能力，上层的历史记忆、上下文工程和 Harness 代码重写都无法有效运作。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="150"><span leaf=""><span textstyle="" style="font-weight: bold;">Harness 与后训练的耦合进化。</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Apodex 的实践揭示了一个前文尚未充分讨论的动力学：</span><span textstyle="" style="font-weight: bold;">Harness 进化与后训练进化难以解耦，二者构成交替推进的耦合关系</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[44]。模型在后训练阶段进化后，Harness 也需要相应进化以匹配新的模型能力；Harness 的改进又反过来为下一轮后训练提供更好的训练环境。这一&#34;两只脚交替踩&#34;的模式与本文 3.5 节 SIA 的&#34;反馈代理决定改 Harness 还是改权重&#34;形成了互补——SIA 将两者视为可选的替代路径，而 Apodex 的实践表明它们在实践中更可能是耦合的协同进化关系。AReaL 2.0 的控制平面（3.6 节）在设计上已隐含了这一耦合——它在记忆、技能、Harness、工具模式和权重之间做选择时，实际上是在不同表面之间推进耦合进化。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="151"><span leaf=""><span textstyle="" style="font-weight: bold;">搜索能力作为自进化的导航器。</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">后训练自进化的核心循环是&#34;诊断自身缺陷 → 针对性造任务和答案 → 训练 → 验证 → 再诊断&#34;。在这一循环中，</span><span textstyle="" style="font-weight: bold;">搜索能力承担着导航器的角色</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">——模型需要搜索与自身弱点对应的资料和数据来构造训练任务 [44]。这使得 Deep Research（深度研究）能力不仅是面向用户的应用，更是自进化循环内部的基础设施。搜索能力本质上代表了&#34;当模型需要提高某个能力时，能否找到对应的数据来提升&#34;这一使能条件。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="152"><span leaf="">3.7.9 递归漂移与多代理验证</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="153"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">&#34;递归漂移&#34;（Recursive Drift）是自进化面临的一个比错误进化更为根本的挑战。Apodex 团队将其精确表述为：</span><span textstyle="" style="font-weight: bold;">模型在自行生成训练数据时，即使最终答案正确，推理过程中的错误也会逐代累积，最终导致进化结果偏离正轨</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[44]。这一问题与本文 4.6 节讨论的错误进化有交集但侧重不同——错误进化关注的是行为策略层面的错误固化，而递归漂移关注的是推理过程层面的误差累积，即使外在表现（答案）正确也可能发生。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="154"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">递归漂移的严重程度与验证难度高度相关。在代码和数学领域，基于规则的验证（单元测试、Lean 形式化验证）能较好地控制漂移，因为验证信号客观且精确。但即便如此，测试代码本身也可能存在&#34;太宽&#34;（让错误代码通过）或&#34;太窄&#34;（惩罚正确代码）的问题，导致细微漂移仍然发生 [44]。在缺乏标准答案的开放性研究领域，递归漂移的控制则几乎缺乏可靠手段。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="155"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Apodex 的应对方案揭示了多代理验证范式的工程价值。其核心设计是</span><span textstyle="" style="font-weight: bold;">Agent Team</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">：问题被分解后，不同子代理分别负责解题和验证，验证代理与解题代理分离以避免上下文污染和角色冲突 [44]。此外引入冗余机制——同一问题由多个独立代理求解，再由全局代理判断哪个答案更可靠。这一设计基于一个结构性的判断：</span><span textstyle="" style="font-weight: bold;">在当前 Self-attention 架构下，上下文越长注意力效果越差，单个模型处理超长上下文存在本质上限</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">，因此多代理协作不是性能优化而是架构必然 [44]。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="156"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一观点为本文 2.3 节的子代理并行模式提供了更深层的理论依据——子代理并行的设计动机不仅是避免上下文污染和提高并行效率，更是为了绕过单模型注意力衰减的结构性约束。Apodex 的实践也表明，裁判（验证器）应在训练过程中一起学习，以避免模型学会讨好裁判的奖励黑客行为 [44]——这与 Self-Harness 的回归测试和 AReaL 2.0 的版本化回滚形成了方法论上的呼应。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="157"><span leaf="">3.7.10 科学品味作为元能力</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="158"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Apodex 团队提出的&#34;发现模型&#34;（Discovery Model）概念指向了自进化的一个更高维度挑战：</span><span textstyle="" style="font-weight: bold;">模型不仅需要解题能力，还需要科学品味——即判断哪些问题值得追求、哪些结果令人惊讶且值得深入的能力</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[44]。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="159"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一观点与 Lilian Weng 在第 4.1 节中提出的&#34;研究品味难以量化&#34;形成互补。Weng 从评估器角度指出研究品味的不可量化性，而 Apodex 的实践则从训练角度揭示了更深层的问题：</span><span textstyle="" style="font-weight: bold;">&#34;会提问&#34;是一种比&#34;会解题&#34;更高维的元能力</span><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">，涉及自我诊断、自我造题和自我训练的后训练循环 [44]。当前模型的训练范式已能较好地培养解题能力（如 SWE-bench Verified 类任务），但如何训练模型提出有价值的创新假设——尤其是 out-of-distribution 的假设——仍是一个未调通的问题。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="160"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这一挑战与本文 2.5.1 节的元方法论转向形成呼应：元方法论要求优化&#34;获取更优答案的机制&#34;，而科学品味恰恰是这一机制中最高阶的组件——它决定了模型在无限的假设空间中应该搜索哪个方向。当前的自进化系统（如 STOP、Self-Harness、DGM）主要在已知任务空间内搜索改进，而科学品味的培养则要求模型具备在未知空间中导航的能力，这可能是实现完全 RSI 的最后一道门槛。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="161"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">关于自进化闭环的实现时间线，Apodex 团队预测最快半年内 AI 即可跑通一次完整的自我进化闭环 [44]。业界数据显示，模型可完成的任务时长约每 7 个月翻一番（从 2024 年 3 月 Claude 3 Opus 的 4 分钟到 2026 年 Claude 4.6 Opus 的 12 小时），这一增速超过了经典摩尔定律 [44]。长程任务能力的指数级增长是 RSI 中&#34;R&#34;（递归）得以实现的前提条件——只有当模型能在无人类监督下持续工作足够长时间时，多轮自我改进才成为可能。</span></span></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="163"><span leaf="">4 未来挑战</span></h1><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="164"><span leaf="">4.1 评估器的脆弱性与模糊性</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="165"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">许多研究声明缺乏快速精确的验证器，现实世界任务亦然。当前自改进循环在评估指标可测量且客观的任务上表现最佳，类似于 RL 的适用条件。研究品味、新颖性和长期科学价值则难以量化——研究品味往往混合了问题构建、实验设计以及对哪些令人惊讶的结果值得追求、哪些失败案例值得重试的判断。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="166"><span leaf="">4.2 上下文与记忆生命周期</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="167"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">随着 AI agent变得更加自主，记忆管理需求持续增长。有用的 Harness 需要管理上下文和记忆，以弥补长上下文生成的现有局限，同时最大化长时程任务的成功率。人类能够在一生中维护记忆，这暗示上下文工程应成为智能的核心组成部分，而非停留在软件系统层。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="168"><span leaf="">4.3 负面结果的利用</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="169"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">文献中成功结果的发表偏差使得 LLM 在判断何时放弃假设、报告负面结果或承认失败方面可能表现不佳。研究 Harness 应使失败尝试易于保留，因为从失败中学习是缩减任务搜索空间的最佳途径。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="170"><span leaf="">4.4 多样性坍缩</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="171"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">进化和 RL 循环倾向于利用已知的高奖励模式。需要防止种群坍缩为同一解决方案的变体，这对于开放式研究尤为关键——最佳路径在当前评估器下可能最初表现更差。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="172"><span leaf="">4.5 奖励黑客</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="173"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">自改进循环优化其获得的任何信号。如果奖励来自单元测试，agent可能过拟合测试；如果来自裁判模型，可能学习针对该裁判的奖励黑客技巧；如果来自基准分数，可能利用基准制品。评估器和权限控制应位于进化 Harness 的循环之外，配合保留测试、轨迹审计和关键决策点的人工审查。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="174"><span leaf="">4.6 错误进化</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="175"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">自进化并非总是朝正确方向演进。agent可能从有限经验中归纳出错误的因果推断，并将其固化为行为策略，导致性能退化而非提升 [41]。这一问题比奖励黑客更为隐蔽——错误进化不一定源于奖励信号被刻意优化，而可能源于agent从非代表性样本中学习。例如，agent在某类任务上反复成功后，可能错误地将成功归因于无关因素，并在新任务中重复不适用策略。这要求自进化系统具备错误检测与回滚机制（如 AReaL 2.0 控制平面中的版本化回滚和 Self-Harness 中的回归测试），而非无条件接受所有&#34;学到的&#34;改进。GDPevo 基准的&#34;规则杂交&#34;设计 [40] 从评测侧对抗了这一问题——只有真正泛化了规则的agent才能通过测试，背答案的agent会被识别。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="176"><span leaf="">4.7 长期成功</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="177"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">以编码agent为例，许多优化目标仍然过于短期。agent可以完成手头任务，但如何保护由数百或数千名工程师共同维护的代码库的长期健康则不够明确。标准沙箱 RLVR 式训练很少涵盖可维护性、所有权边界、迁移成本、向后兼容性或未来调试负担。</span></span></p><h2 style="color: #2B77BF;text-align: start;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="178"><span leaf="">4.8 人类的角色</span></h2><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="179"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">人类应向栈的上层移动而非被移出循环，即在正确的抽象层级和正确的时间点提供监督。系统设计应考虑何时以及如何设置此类接触点。上述诸多挑战需要人类的反馈和引导——毕竟，我们构建技术是为了人类更美好的未来，而非相反。</span></span></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="181"><span leaf="">5 结论</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="182"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Harness 工程正从提示模板组合演变为运行时系统设计，其在 RSI 路径中的地位日益关键。本文梳理了三大设计模式（工作流自动化、文件系统持久记忆、子agent并行）、六条优化路径（上下文工程、工作流设计、自我改进循环、进化搜索、权重联合优化、面向自演进的系统基础设施）以及自演进agent的理论框架与生态全景（递归自指、轻量化范式、多agent协同、标准化协议、评测基准、错误进化风险、Skill 分层与产业实践、使能条件与耦合动力学、递归漂移与多代理验证、科学品味作为元能力），揭示了从手工启发式到自动化搜索再到系统级治理的方法论迁移趋势。</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="183"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">研究发现，一旦 Harness 设计成为可执行的搜索空间，强大的编码agent能够利用与人类工程师相同的设计空间进行探索。AReaL 2.0 的工作进一步表明，自演进agent在企业级部署中的瓶颈不在 RL 算法，而在系统基础设施——需要标准化的轨迹数据协议、企业级数据agent和统一的演进控制平面三大支柱协同。与此同时，Gödel Agent 和 DGM 从递归自指与开放式进化两个方向推进了理论边界，Reflexion 和 MetaAgent 证明了轻量化自进化范式的可行性，GDPevo 基准表明当前agent已具备一定的自进化能力（准确率提升 17–22%）。崔添翼提出的&#34;Skill 是 Harness 自进化的初级形式&#34;为优化路径提供了更精细的梯度划分，而头部 AI 公司向&#34;模型 + Harness&#34;双轮驱动模式的迁移则标志着产业级共识的形成。Apodex 的工业实践进一步揭示了编码能力作为自进化底座、Harness 与后训练的耦合进化、搜索能力作为自进化导航器等使能条件，以及递归漂移和多代理验证、科学品味作为元能力等深层挑战。然而，递归结构本身不足以驱动自我改进——基座模型必须具备足够的智能才能改进机制。评估器质量、多样性维持、奖励黑客、错误进化、递归漂移等问题仍是实现完全 RSI 的关键瓶颈。未来研究需要在自动化效率与人类监督之间找到平衡，使 Harness 改进既可扩展又可控。</span></span></p><hr style="box-sizing: content-box;height: 2px;margin: 16px 0px;border: 0px none;background-color: rgba(0, 0, 0, 0.9);overflow: hidden;padding: 0px;color: #2B77BF;"/><h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="185"><span leaf="">参考文献</span></h1><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="186"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[1] Good, I. J. &#34;Speculations Concerning the First Ultraintelligent Machine.&#34;Advances in Computers, 6:31–88, 1965.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="187"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[2] Yudkowsky, Eliezer. &#34;Recursive Self-Improvement.&#34; LessWrong, 2008.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="188"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[3] Anthropic. &#34;Recursive Self-Improvement.&#34; Anthropic Institute, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="189"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[4] OpenAI. &#34;How Agents Are Transforming Work.&#34; OpenAI Blog, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="190"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[5] Karpathy, A. &#34;autoresearch.&#34; GitHub repository, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="191"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[6] Weng, Lilian. &#34;Prompt Engineering.&#34; Lil&#39;Log, Mar 2023.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="192"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[7] Zhang, et al. &#34;Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.&#34; ICLR 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="193"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[8] Ye, et al. &#34;Meta Context Engineering via Agentic Skill Evolution.&#34; arXiv:2601.21557, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="194"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[9] Lee, et al. &#34;Meta-Harness: End-to-End Optimization of Model Harnesses.&#34; arXiv:2603.28052, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="195"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[10] Lu, et al. &#34;Towards End-to-End Automation of AI Research.&#34;Nature, 651:914–919, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="196"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[11] Meng, et al. &#34;ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence.&#34; arXiv:2605.26340, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="197"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[12] Kulikov, et al. &#34;Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data.&#34; arXiv:2606.25996, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="198"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[13] Hu, Lu, and Clune. &#34;Automated Design of Agentic Systems.&#34; ICLR 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="199"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[14] Zhang, et al. &#34;AFlow: Automating Agentic Workflow Generation.&#34; ICLR 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="200"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[15] Zelikman, et al. &#34;Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation.&#34; COLM 2024.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="201"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[16] Zhang, et al. &#34;Self-Harness: Harnesses That Improve Themselves.&#34; arXiv:2606.09498, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="202"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[17] Novikov, et al. &#34;AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery.&#34; arXiv:2506.13131, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="203"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[18] Zhang, et al. &#34;Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents.&#34; arXiv:2505.22954, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="204"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[19] Hebbar, et al. &#34;SIA: Self Improving AI with Harness &amp; Weight Updates.&#34; arXiv:2605.27276, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="205"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[20] Trehan and Chopra. &#34;Why LLMs Aren&#39;t Scientists Yet: Lessons from Four Autonomous Research Attempts.&#34; arXiv:2601.03315, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="206"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[21] Fernando, et al. &#34;Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution.&#34; arXiv:2309.16797, 2023.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="207"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[22] Agrawal, A. et al. &#34;GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.&#34; arXiv:2507.19457, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="208"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[23] Lange, Imajuku, and Cetin. &#34;ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution.&#34; arXiv:2509.19349, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="209"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[24] Wang, et al. &#34;ThetaEvolve: Test-time Learning on Open Problems.&#34; arXiv:2511.23473, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="210"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[25] Zhang, et al. &#34;Hyperagents.&#34; arXiv:2603.19461, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="211"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[26] Madaan, et al. &#34;Self-Refine: Iterative Refinement with Self-Feedback.&#34; NeurIPS 2023.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="212"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[27] Bubeck, et al. &#34;Early Science Acceleration Experiments with GPT-5.&#34; arXiv:2511.16072, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="213"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[28] Yan, et al. &#34;Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents.&#34; arXiv:2607.01120, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="214"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[29] AReaL Project. &#34;AReaL 2.0: Agent-Oriented Online RL Infrastructure.&#34; GitHub repository, <a href="https://github.com/areal-project/AReaL," target="_blank">https://github.com/areal-project/AReaL,</a> 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="215"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[30] Gao, et al. &#34;A Survey of Self-Evolving Agents: On Path to Artificial Super Intelligence.&#34; arXiv:2507.21046, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="216"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[31] Fang, et al. &#34;A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.&#34; arXiv:2508.07407, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="217"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[32] Yin, et al. &#34;Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement.&#34;Proceedings of ACL, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="218"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[33] &#34;Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair.&#34; arXiv, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="219"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[34] Shinn, et al. &#34;Reflexion: Language Agents with Verbal Reinforcement Learning.&#34;NeurIPS, 2023.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="220"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[35] Qian and Liu. &#34;MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning.&#34; arXiv:2508.00271, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="221"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[36] Xia, et al. &#34;Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning.&#34;ICLR 2026 Workshop RSI, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="222"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[37] Lu, et al. &#34;MorphAgent: Empowering Agents through Self-Evolving Profiles and Decentralized Collaboration.&#34; arXiv:2410.15048, 2025.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="223"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[38] Zhang, et al. &#34;Autogenesis: A Self-Evolving Agent Protocol.&#34; arXiv:2604.15034, 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="224"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[39] Xi, et al. &#34;AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.&#34; arXiv:2406.04151, 2024.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="225"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[40] Prism-Shadow. &#34;GDPevo: A Benchmark for Evaluating Agent Self-Evolution on GDP-Related Tasks.&#34; 2026. <a href="https://github.com/Prism-Shadow/GDPevo" target="_blank">https://github.com/Prism-Shadow/GDPevo</a></span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="226"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[41] Ren, et al. &#34;Misevolution: On the Risk of Error Evolution in Self-Evolving Agents.&#34; 2026.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="227"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[42] 崔添翼（DeepSeek）. 社交媒体转发评论. <a href="https://x.com/tianyi/status/2074475185957380379," target="_blank">https://x.com/tianyi/status/2074475185957380379,</a> 2026年7月.</span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="228"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[43] 量子位. &#34;翁荔新博客提出「自进化先从Harness开始」，DeepSeek崔添翼转发附议.&#34; 2026年7月. <a href="https://www.qbitai.com/2026/07/442134.html" target="_blank">https://www.qbitai.com/2026/07/442134.html</a></span></span></p><p style="text-align: start;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="229"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">[44] 硅谷101播客. &#34;E242｜最快半年AI跑通自进化？与陈天桥首席科学家聊聊硅谷模型必争之地.&#34; 2026年7月. 嘉宾：杜少雷（Apodex首席科学家、华盛顿大学副教授）、Beibin Li（Apodex首席科学家）. <a href="https://www.youtube.com/watch?v=VGIbhIW5ljk" target="_blank">https://www.youtube.com/watch?v=VGIbhIW5ljk</a></span></span></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=0b6d4a81&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486117%26idx%3D1%26sn%3D619c8f86e5c89b279f9de19a8dd72910">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sat, 11 Jul 2026 09:59:00 +0800</pubDate>
    </item>
    <item>
      <title>AI时代下对安全攻防演进的思考（上篇）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486092&amp;idx=1&amp;sn=f2966b6cadd0258e85b69ffe06c7dee8</link>
      <description>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</description>
      <content:encoded><![CDATA[<p><span>riusksk</span> <span>2026-07-09 18:02</span> <span style="display: inline-block;">广东</span></p>




  <p>以下文章来源于：华为安全应急响应中心</p>
  <strong>华为安全应急响应中心</strong>
  <p>华为安全应急响应中心（HUAWEI PSIRT）官方公众号。</p>



  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=c30ef9d5&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FEH8ujwq9pxiaYkHLFRDVS0keuCnicZD9X5eQhJa6bHS01Tlw0oAicHicrsf94mBFEhDLkLakBZiaVjybP3RAKJibDibxnzqLCqKsu6ich1MwWrJYpno%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</p>
  <div style="box-sizing: border-box;font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);" data-pm-slice="0 0 []"><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">引言</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;text-indent: 2em;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><span leaf="">过去相当长的一段时间里，安全攻防工作的基本现实是：高质量漏洞挖掘、稳定漏洞利用与复杂逆向分析，都高度依赖少数经验密集型研究员。无论是大规模代码理解、崩溃归因、利用链设计，还是样本行为复原和保护壳突破，背后都需要较高的人工投入与长期训练。因此，企业的漏洞响应节奏、安全团队分工乃至外包服务定价，实际上都默认建立在“高级研究能力稀缺”这一前提之上。</span></p><span style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);box-sizing: border-box;" data-pm-slice="0 0 []"><span leaf="">     这一前提正在松动。</span><span lang="EN-US"><span leaf="">2024</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">6</span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Google Project Zero </span></span><span leaf="">发布</span><span lang="EN-US"><span leaf=""> Project Naptime</span></span><span leaf="">，公开展示了大模型结合代码浏览、调试、验证工具后，在漏洞研究任务中的潜力</span><sup><span lang="EN-US"><span leaf="">[1]</span></span></sup><span leaf="">。同年</span><span lang="EN-US"><span leaf="">11</span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Big Sleep </span></span><span leaf="">进一步披露其在</span><span lang="EN-US"><span leaf=""> SQLite </span></span><span leaf="">中发现了此前未知、具有利用价值的现实世界内存安全漏洞</span><sup><span lang="EN-US"><span leaf="">[2]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">2025</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">7</span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Google </span></span><span leaf="">又宣布相关系统已能结合威胁情报，在漏洞进入在野利用链之前完成前置识别</span><sup><span lang="EN-US"><span leaf="">[3]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">2026</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">4</span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Anthropic </span></span><span leaf="">通过</span><span lang="EN-US"><span leaf=""> Project Glasswing </span></span><span leaf="">与</span><span lang="EN-US"><span leaf=""> Claude Mythos Preview </span></span><span leaf="">把同一问题推进到更敏感的位置：模型不只是作为实验对象被评测，而是开始被有选择地引入关键软件和基础设施的防御流程</span><sup><span lang="EN-US"><span leaf="">[4][5]</span></span></sup><span leaf="">。</span></span><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><span leaf="">如果这些进展只是个别样例，它们的行业意义有限；但当类似能力同时出现在真实代码库、厂商防御合作、国家级竞赛与公开 benchmark 中时，就已经不宜再被视作零散演示。真正值得追问的，其实有三个问题：</span></p></div><ul style="list-style-type: disc;box-sizing: border-box;padding-left: 20px;list-style-position: outside;" class="list-paddingleft-1"><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;"><span leaf="">AI 到底改变Web 安全、Pwn 与逆向工程中的哪些工作环节（因为web/pwn/reverse是安全攻防三大核心领域，AI对安全攻防的影响必然会优先在三大领域中体现）？</span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;"><span leaf="">这种变化对企业、安全人员与行业发展意味着什么？</span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;"><span leaf="">未来三到五年，安全攻防的重心又会沿着什么方向迁移？</span></p></li></ul><div style="box-sizing: border-box;font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);" data-pm-slice="0 0 []"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span data-pm-slice="0 0 []"><span leaf="" data-pm-slice="1 1 [&#34;para&#34;,{&#34;tagName&#34;:&#34;section&#34;,&#34;attributes&#34;:{&#34;style&#34;:&#34;box-sizing: border-box;font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);&#34;,&#34;data-pm-slice&#34;:&#34;0 0 []&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;},&#34;list&#34;,{&#34;type&#34;:&#34;ul&#34;,&#34;style&#34;:&#34;list-style-type: disc;box-sizing: border-box;padding-left: 20px;list-style-position: outside;&#34;,&#34;class&#34;:&#34;list-paddingleft-1&#34;,&#34;start&#34;:null},&#34;listitem&#34;,{&#34;style&#34;:&#34;box-sizing: border-box;&#34;},&#34;para&#34;,{&#34;tagName&#34;:&#34;p&#34;,&#34;attributes&#34;:{&#34;style&#34;:&#34;white-space: normal; margin: 0px; padding: 0px; box-sizing: border-box; text-indent: 2em;&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;},&#34;node&#34;,{&#34;tagName&#34;:&#34;span&#34;,&#34;attributes&#34;:{&#34;style&#34;:null,&#34;data-pm-slice&#34;:&#34;0 0 []&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;}]"><span textstyle="" style="color: rgb(122, 68, 66);font-weight: normal;">本文是《AI时代下对安全攻防演进的思考》系列文章的第一篇，本系列将分为三篇系统阐述AI技术对安全攻防领域带来的深刻变革。本篇将重点探讨AI在安全攻防领域的现状与基础能力，后续两篇将分别深入分析三大核心领域的结构性重塑以及深远影响与未来展望。</span></span></span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">从汽车、织机到 ATM：技术革命如何改写安全攻防</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><span style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">      汽车刚出现时，问题并不只是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">谁会开车</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是整个道路系统都没准备好。美国联邦公路管理局（</span><span lang="EN-US"><span leaf="">Federal Highway Administration</span></span><span leaf="">，</span><span lang="EN-US"><span leaf="">FHWA</span></span><span leaf="">）回顾过那段历史：二十世纪初，美国不少道路上的指示牌由不同汽车俱乐部分散设置，在部分主要路线中，单一路线甚至可能同时出现</span><span lang="EN-US"><span leaf="">11</span></span><span leaf="">种不同标识；之后才逐步形成中心线、红绿灯、停车标志以及统一交通控制手册（</span><span lang="EN-US"><span leaf="">MUTCD</span></span><span leaf="">）这样的标准体系</span><sup><span lang="EN-US"><span leaf="">[7]</span></span></sup><span leaf="">。这段历史的关键不在汽车本身，而在它说明了一件事：</span><span textstyle=""><b><span leaf="">一项工具能力越过阈值之后，最先暴露出来的，往往不是个人是否足够努力，而是接口不统一、规则不清晰、责任不可追溯、基础设施跟不上。</span></b></span></span><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">这也是今天讨论 AI for 安全攻防时最容易被忽略的地方。围绕 AI 的讨论，常见的两种偏差恰好相反：一种把少数高光案例直接外推成全面替代，另一种则把部分任务自动化直接理解成职业失去价值。两种看法都太急着讨论“工具会不会替代人”，却没有关注到历史长河中相似事件所反映出的关键现象：<span textstyle="" style="font-weight: bold;">历史上真正改变行业格局的，往往不是某个新工具本身，而是围绕它发生的一整套流程、标准、分工和责任重写。</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">职业结构的变化也遵循同样的逻辑。美国劳工统计局（</span><span lang="EN-US"><span leaf="">Bureau of Labor Statistics</span></span><span leaf="">，</span><span lang="EN-US"><span leaf="">BLS</span></span><span leaf="">）关于职业变迁的研究，以及其后对</span><span lang="EN-US"><span leaf=""> 1860 </span></span><span leaf="">年到</span><span lang="EN-US"><span leaf=""> 2015 </span></span><span leaf="">年美国职业结构的回顾，都指出技术扩散带来的不是岗位线性消失，而是职业构成和任务重心的持续迁移</span><sup><span lang="EN-US"><span leaf="">[6][25]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">BLS </span></span><span leaf="">的表述很直接：企业一旦采用新技术、开发新产品或改变经营方式，职业结构也会跟着变化</span><sup><span lang="EN-US"><span leaf="">[25]</span></span></sup><span leaf="">。换句话说，技术革命真正改写的，通常不是岗位名字，而是岗位内部的任务组合、所需技能，以及与其他岗位的协作关系。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">工业革命和信息革命里的很多例子都能说明这一点。十九世纪动力织机把每码布所需劳动量压缩了</span><span lang="EN-US"><span leaf=""> 98%</span></span><span leaf="">，但织工岗位并没有因此简单消失；随着布匹价格下降、需求扩大，以及多台织机协同管理等新技能的重要性上升，织工岗位反而增加</span><sup><span lang="EN-US"><span leaf="">[27]</span></span></sup><span leaf="">。自动取款机普及之后，银行柜员的人均现金处理工作减少了，但岗位并没有像很多人预想的那样被整体抹掉；柜员反而从单纯办业务转向关系维护、产品销售和更复杂的人机协作任务</span><sup><span lang="EN-US"><span leaf="">[27]</span></span></sup><span leaf="">。这就是</span><span lang="EN-US"><span leaf=""> “</span></span><span leaf="">杰文斯悖论</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">：在技术进步提高资源利用效率时，该资源的使用成本下降，反而刺激需求大幅扩张，导致资源总消耗量不降反升的现象。历史反复提醒我们的，不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">技术不会替代工作</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">技术很少按最直线的方式替代工作</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。它更常见的路径，是先压缩一部分重复性任务，再抬高剩余任务对理解、协调和判断的要求。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">还有一层常被低估的变化：通用技术的价值，往往要靠配套重构才能兑现。美国国家经济研究局（</span><span lang="EN-US"><span leaf="">National Bureau of Economic Research</span></span><span leaf="">，</span><span lang="EN-US"><span leaf="">NBER</span></span><span leaf="">）关于美国制造业电气化的研究发现，电力带来的生产率提升并不是孤立发生的，而是迅速伴随着资本深化和组织结构调整</span><sup><span lang="EN-US"><span leaf="">[26]</span></span></sup><span leaf="">。另一篇关于通用技术的经典研究则指出，蒸汽机、电动机、半导体和计算机这类技术的收益不会自动落地，它们依赖下游行业的配套创新和组织重构；如果两端之间只有松散的市场交易关系，创新往往会来得太慢、太晚</span><sup><span lang="EN-US"><span leaf="">[28]</span></span></sup><span leaf="">。</span><b><span leaf="">放到今天，单纯问</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">模型强不强</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">其实是不够的，更关键的问题是：组织有没有准备好围绕这种能力重写流程。</span></b></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">沿着这条历史线再看今天的安全攻防，很多变化就没有那么难理解了。AI 带来的不会只是“多一个更强的扫描器”或“少几个初级岗位”，而是整个安全生产方式的重排：哪些漏洞会先失去稀缺性，哪些验证与修补环节会成为瓶颈，哪些工具会从“给人用”转成“给 AI 调用”，以及企业为什么必须重写授权、审计、回滚和责任链条。也正因为如此，后文才需要分别讨论 Web 安全、Pwn 与逆向工程：这三类方向代表了三种很不一样的任务形态，也最能看清 AI 会先在哪里形成规模收益，又会在哪里被复杂性和对抗性重新拉高门槛。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">AI 自动攻防的现实进展：从能力演示到工程对象</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">在安全领域中最关键的变化，不是模型回答安全问题时听起来更像人类专家，而是它开始在工具链与交互环境的支持下，成为能够持续推进任务的工程对象。</span><span lang="EN-US"><span leaf="">Project Naptime </span></span><span leaf="">的意义正体现在这里。</span><span lang="EN-US"><span leaf="">Google </span></span><span leaf="">并未把模型当成一个孤立问答器，而是把它放进带有代码检索、调试、验证和多轨采样的环境中</span><sup><span lang="EN-US"><span leaf="">[1]</span></span></sup><span leaf="">。这一点十分重要，因为安全研究从来不是一次性求解，而是围绕假设提出、证据收集、结果验证和错误修正的反复循环。模型一旦具备这类工作脚手架，它的能力表现就不再主要取决于</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">会不会答对</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而更多取决于能否在反馈中持续推进。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span lang="EN-US"><span leaf="">Big Sleep </span></span><span leaf="">则把这一能力第一次较清楚地拉进现实软件。</span><span lang="EN-US"><span leaf="">Google 2024</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">11</span></span><span leaf="">月披露的</span><span lang="EN-US"><span leaf=""> SQLite </span></span><span leaf="">案例，并不是对已知漏洞的机械复述，而是围绕真实代码变更做进一步挖掘，最终发现此前未知、具有利用价值的缺陷</span><sup><span lang="EN-US"><span leaf="">[2]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">2025</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">7</span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Google </span></span><span leaf="">又宣称其系统已能结合威胁情报，对即将进入在野利用阶段的漏洞做前置识别</span><sup><span lang="EN-US"><span leaf="">[3]</span></span></sup><span leaf="">。无论这些案例未来是否还能大规模复现，它们已经说明：</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">参与漏洞发现的讨论，至少应当从</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">能不能做演示</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">转向</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">能否稳定进入现实流程</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span lang="EN-US"><span leaf="">Anthropic </span></span><span leaf="">的一系列披露则把问题进一步推进到</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">能力治理</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">的层面。</span><span lang="EN-US"><span leaf="">Project Glasswing </span></span><span leaf="">的核心不只是展示模型能力，而是讨论更强的挖洞能力是否应当优先部署到关键软件与基础设施防御中</span><sup><span lang="EN-US"><span leaf="">[4]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">Claude Mythos Preview </span></span><span leaf="">作为配套模型，则在厂商自评框架下展示了模型在部分</span><span lang="EN-US"><span leaf="">0day</span></span><span leaf="">挖掘与利用任务上的能力跃迁</span><sup><span lang="EN-US"><span leaf="">[5]</span></span></sup><span leaf="">。与此同时，</span><span lang="EN-US"><span leaf="">Anthropic </span></span><span leaf="">还专门发布了</span><span lang="EN-US"><span leaf="">AI</span></span><span leaf="">发现漏洞的协调披露原则，说明其已将模型能力视作可能影响传统</span><span lang="EN-US"><span leaf=""> CVD</span></span><span leaf="">（</span><span lang="EN-US"><span leaf="">Coordinated Vulnerability Disclosure</span></span><span leaf="">，协调漏洞披露）机制的现实变量</span><sup><span lang="EN-US"><span leaf="">[20]</span></span></sup><span leaf="">。这些披露当然带有明确的厂商立场，但若将其与</span><span lang="EN-US"><span leaf=""> Google </span></span><span leaf="">的路线、</span><span lang="EN-US"><span leaf="">DARPA </span></span><span leaf="">的长期计划以及竞赛评测的发展放在一起观察，其共同指向已经相当明确：</span><b><span lang="EN-US"><span leaf="">AI</span></span><span leaf="">在安全中的角色，正在从</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">辅助说明问题</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">走向</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">参与生产问题发现结果</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span></b></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">竞赛与</span><span lang="EN-US"><span leaf=""> Benchmark </span></span><span leaf="">为这种判断提供了更稳的外部支撑。</span><span lang="EN-US"><span leaf="">DARPA </span></span><span leaf="">从</span><span lang="EN-US"><span leaf=""> Cyber Grand Challenge </span></span><span leaf="">到</span><span lang="EN-US"><span leaf=""> AIxCC</span></span><span leaf="">，持续十余年推动自动发现与自动修补的路线</span><sup><span lang="EN-US"><span leaf="">[9][10]</span></span></sup><span leaf="">；</span><span lang="EN-US"><span leaf="">Anthropic 2025 </span></span><span leaf="">年关于</span><span lang="EN-US"><span leaf=""> Cyber Competitions </span></span><span leaf="">的披露则表明，模型在中等难度竞赛中已足以接近不少人类选手，但在最高难度题目上仍暴露明显短板</span><sup><span lang="EN-US"><span leaf="">[11]</span></span></sup><span leaf="">。国内的腾讯云智能渗透挑战赛，则进一步显示多智能体渗透、自动利用与评测框架已成为可组织、可复现的工程方向</span><sup><span lang="EN-US"><span leaf="">[8][24]</span></span></sup><span leaf="">，赛制也从单点解题推进到双线并行的更接近真实攻防的模式：一条赛线要求智能体围绕信息收集、漏洞发现、利用执行和权限维持完成长链路渗透；另一条赛线则把智能体放入信息不对称、多主体互动和提示词对抗环境中，考察其协作、博弈与上下文管理能力</span><sup><span lang="EN-US"><span leaf="">[24]</span></span></sup><span leaf="">。环境设计同时引入真实</span><span lang="EN-US"><span leaf=""> CVE</span></span><span leaf="">、云安全缺陷和</span><span lang="EN-US"><span leaf=""> AD </span></span><span leaf="">域渗透拓扑，这一点尤其重要，因为它让比赛不再只是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">能不能答出一道题</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而更接近</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">能不能在复杂环境里把任务推进下去</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">更重要的是，这类比赛带来的价值并不只是展示</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">会什么，更在于暴露</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">还不会什么。从外部对比赛的复盘信息看，和首届相比，智能体在半年内已经从单一漏洞利用迈向全流程渗透尝试，整体表现接近初级渗透测试人员；但与此同时，复杂环境下的控场能力、长链路推理、路径回溯、错误恢复、记忆管理和环境清理，依旧是反复暴露出来的短板</span><sup><span lang="EN-US"><span leaf="">[24]</span></span></sup><span leaf="">。这意味着当前</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">在攻防中的主要问题，并不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">完全做不到</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是经常在局部环节看似可用，却难以把完整攻击链稳定打通。也正因此，</span><b><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">安全能力今天所处的位置并不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">已经统治高难度攻防</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">已经足以改写大量中低难度任务的生产方式，同时也在真实场景中暴露出新的系统性瓶颈</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span></b></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">小结</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span data-pm-slice="0 0 []"><span leaf="">本文回顾了AI在安全攻防领域从概念演示迈向工程应用的关键进展。这些能力具体将如何重塑Web安全、Pwn与逆向工程这三大核心战场？基础漏洞是否会批量消失？攻防的稀缺能力又将向何处迁移？我们将在《AI时代下对安全攻防演进的思考（中篇）》深入解析各领域面临的挑战与机遇。</span></span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">参考文献</span></strong></p></div></div></div></div></div><div style="text-align: center;box-sizing: border-box;"><div style="display: inline-block;width: 100%;height: 240px;vertical-align: top;overflow-y: auto;box-sizing: border-box;"><p style="font-size: 12px;text-align: left;box-sizing: border-box;"><ol style="list-style-type: decimal;box-sizing: border-box;padding-left: 20px;list-style-position: outside;" class="list-paddingleft-2"><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，Project Naptime: Evaluating      Offensive Security Capabilities of Large Language Models，2024-06-20。<a href="https://projectzero.google/2024/06/project-naptime.html" target="_blank">https://projectzero.google/2024/06/project-naptime.html</a></span></p></li><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，From Naptime to Big Sleep: Using      Large Language Models To Catch Vulnerabilities In Real-World Code，2024-11-01。<a href="https://projectzero.google/2024/10/from-naptime-to-big-sleep.html" target="_blank">https://projectzero.google/2024/10/from-naptime-to-big-sleep.html</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google，A summer of security: empowering cyber defenders      with AI，2025-07-15。<a href="https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/" target="_blank">https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Project Glasswing: Securing critical software for      the AI era，2026-04-07。<a href="https://www.anthropic.com/glasswing" target="_blank">https://www.anthropic.com/glasswing</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude Mythos Preview，2026-04-07。<a href="https://red.anthropic.com/2026/mythos-preview/" target="_blank">https://red.anthropic.com/2026/mythos-preview/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Occupational changes during the      20th century，2006-03。<a href="https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm" target="_blank">https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Federal Highway Administration，The Evolution of MUTCD。<a href="https://mutcd.fhwa.dot.gov/kno-history.htm" target="_blank">https://mutcd.fhwa.dot.gov/kno-history.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">腾讯云黑客松官网；腾讯云开发者社区《400+极客菁英共聚羊城，见证国内首个AI智能渗透挑战赛》。<a href="https://tch.cloud.tencent.com/" target="_blank">https://tch.cloud.tencent.com/</a> ；<a href="https://cloud.tencent.com/developer/article/2651925" target="_blank">https://cloud.tencent.com/developer/article/2651925</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，AIxCC:      AI Cyber Challenge；AI Cyber Challenge marks      pivotal inflection point for cyber defense。<a href="https://www.darpa.mil/research/programs/ai-cyber" target="_blank">https://www.darpa.mil/research/programs/ai-cyber</a> ；<a href="https://www.darpa.mil/news/2025/aixcc-results" target="_blank">https://www.darpa.mil/news/2025/aixcc-results</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，Cyber      Grand Challenge (CGC)。<a href="https://www.darpa.mil/research/programs/cyber-grand-challenge" target="_blank">https://www.darpa.mil/research/programs/cyber-grand-challenge</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude is competitive with humans      in (some) cyber competitions，2025-08-09。<a href="https://red.anthropic.com/2025/cyber-competitions/" target="_blank">https://red.anthropic.com/2025/cyber-competitions/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">National Institute of Standards and      Technology (NIST)，Artificial      Intelligence Risk Management Framework: Generative Artificial Intelligence      Profile，2024-07-26。<a href="https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence" target="_blank">https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Open Worldwide Application Security      Project (OWASP) GenAI Security Project，OWASP Top 10 for LLM is now the GenAI Security      Project and promoted to OWASP Flagship status，2025-03-26。<a href="https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/" target="_blank">https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Partnering with Mozilla to      improve Firefox&#39;s security，2026-03-06。<a href="https://red.anthropic.com/2026/firefox/" target="_blank">https://red.anthropic.com/2026/firefox/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Hanzhuo Tan, Qi Luo, Jing Li, Yuqun      Zhang，LLM4Decompile:      Decompiling Binary Code with Large Language Models，EMNLP 2024。<a href="https://aclanthology.org/2024.emnlp-main.203/" target="_blank">https://aclanthology.org/2024.emnlp-main.203/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anton Tkachenko, Dmitrij Suskevic,      Benjamin Adolphi，Deconstructing      Obfuscation: A four-dimensional framework for evaluating Large Language      Models assembly code deobfuscation capabilities，2025。<a href="https://arxiv.org/abs/2505.19887" target="_blank">https://arxiv.org/abs/2505.19887</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Abraham Clements et al.，Towards LLM-Resistant Software      Protection: Agent Failure Patterns in CTF Reverse Engineering，NDSS BAR 2026。<a href="https://www.ndss-symposium.org/ndss-paper/auto-draft-657/" target="_blank">https://www.ndss-symposium.org/ndss-paper/auto-draft-657/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Cybench: A Framework for Evaluating      Cybersecurity Capabilities and Risks of Language Models。<a href="https://cybench.github.io/" target="_blank">https://cybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">BountyBench: Dollar Impact of AI Agent      Attackers and Defenders on Real-World Cybersecurity Systems。<a href="https://bountybench.github.io/" target="_blank">https://bountybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Coordinated vulnerability disclosure for      Claude-discovered vulnerabilities，最后更新于 2026-03-06。<a href="https://www.anthropic.com/coordinated-vulnerability-disclosure" target="_blank">https://www.anthropic.com/coordinated-vulnerability-disclosure</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Node.js Project，New HackerOne Signal Requirement      for Vulnerability Reports，最后更新于 2026-02-19。<a href="https://nodejs.org/en/blog/announcements/hackerone-signal-requirement" target="_blank">https://nodejs.org/en/blog/announcements/hackerone-signal-requirement</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">HackerOne，Code of Conduct。<a href="https://www.hackerone.com/policies/code-of-conduct" target="_blank">https://www.hackerone.com/policies/code-of-conduct</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Daniel Stenberg，The end of the curl bug-bounty，2026-01-26。<a href="https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/" target="_blank">https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">前沿观察 赛事纪实：从腾讯云黑客松，洞见智能体时代的攻防新格局，2026-04-17。<a href="https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q" target="_blank">https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Changes in the U.S. occupational      mix from 1860 to 2015，2019-08。<a href="https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm" target="_blank">https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Martin Fiszbein, Jeanne Lafortune, Ethan      G. Lewis, José Tessada，Powering Up Productivity: The Effects of Electrification on      U.S. Manufacturing，National Bureau of Economic      Research (NBER) Working Paper 28076，2020；2024-04 修订。<a href="https://www.nber.org/papers/w28076" target="_blank">https://www.nber.org/papers/w28076</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">James Bessen，Toil and Technology，International Monetary Fund (IMF) Finance &amp; Development，2015-03。<a href="https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm" target="_blank">https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Timothy F. Bresnahan, Manuel Trajtenberg，General Purpose Technologies      &#34;Engines of Growth?&#34;，National Bureau of      Economic Research (NBER) Working Paper 4148，1992-08。<a href="https://www.nber.org/papers/w4148" target="_blank">https://www.nber.org/papers/w4148</a></span></p></li></ol></p></div></div></div><p style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf=""><span textstyle="" style="font-style: italic;">系列连载未完待续，下篇敬请期待。</span></span></p><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box;overflow-wrap: break-word !important;clear: both;min-height: 1em;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;" data-pm-slice="0 0 []"><strong data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;letter-spacing: 0.544px;display: inline;color: rgb(62, 62, 62);font-family: 楷体;text-align: left;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;letter-spacing: normal;color: rgb(255, 0, 0);font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><strong style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;display: inline;color: rgb(255, 79, 121);font-size: 16px;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">本公众号发布、转载的文章所涉及的技术、思路、工具仅供学习交流，任何人不得将其用于非法用途及盈利等目的，否则后果自行承担！</span></span></strong></span></strong></p><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;text-align: center;"><span data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">点这里 <img alt="图片" class="rich_pages wxw-img __bg_gif" data-aistatus="1" data-imgfileid="100041710" data-ratio="0.5982532751091703" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;vertical-align: middle;height: auto !important;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;font-family: PingFangSC-Regular, &#34;PingFang SC&#34;;text-indent: 28px;color: rgb(62, 62, 62);font-size: 16px;width: 61.9922px !important;visibility: visible !important;" data-w="458" data-width="100%" src="https://wechat2rss.xlab.app/img-proxy/?k=6849f525&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_gif%2FMfTd6rd9CyvNRMW8I9cvI1CK5gKiaYqg2veTn9t9dAe1GxYic7pAvgvRIKNFickConFyX8AvW2reAq8GchJI6aBpA%2F640%3Fwx_fmt%3Dgif%26wxfrom%3D5%26wx_lazy%3D1%26tp%3Dwebp%23imgIndex%3D14"/></span><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" data-pm-slice="1 1 [&#34;para&#34;,null,&#34;node&#34;,{&#34;tagName&#34;:&#34;span&#34;,&#34;attributes&#34;:{&#34;style&#34;:&#34;color: rgb(127, 229, 230); font-family: \&#34;Helvetica Neue\&#34;, Helvetica, \&#34;Hiragino Sans GB\&#34;, \&#34;Microsoft YaHei\&#34;, Arial, sans-serif; font-size: 14px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: 0.578px; orphans: 2; text-align: center; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px;  background-color: rgb(255, 255, 255); text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; display: inline !important; float: none;&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;}]" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">关注我们，一键三连～</span></span></span></p><p class="mp_profile_iframe_wrp" nodeleaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;"><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe js_wx_tap_highlight" data-pluginname="mpprofile" data-nickname="华为安全应急响应中心" data-alias="HUAWEI_PSIRT" data-index="0" data-from="2" data-headimg="http://mmbiz.qpic.cn/sz_mmbiz_png/Pf9eicDVDMxHbPW1POGs9HHQCUGXBXg7u6TCtI2ab5DdIEfxJWcR46krXgudVuibfibsqRYlAtN2RLdaiaOCosQMSw/300?wx_fmt=png&amp;wxfrom=19" data-signature="华为安全应急响应中心（HUAWEI PSIRT）官方公众号。" data-id="MzI0MTY5NDQyMw==" data-is_biz_ban="0" data-origin_num="49" data-biz_account_status="0" data-service_type="1" data-verify_status="2"></mp-common-profile></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=890733f0&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486092%26idx%3D1%26sn%3Df2966b6cadd0258e85b69ffe06c7dee8">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 09 Jul 2026 18:02:00 +0800</pubDate>
    </item>
    <item>
      <title>AI时代下对安全攻防演进的思考（中篇）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486092&amp;idx=2&amp;sn=c494c0340d1a05ef458bec7584c994bf</link>
      <description>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</description>
      <content:encoded><![CDATA[<p><span>riusksk</span> <span>2026-07-09 18:02</span> <span style="display: inline-block;">广东</span></p>




  <p>以下文章来源于：华为安全应急响应中心</p>
  <strong>华为安全应急响应中心</strong>
  <p>华为安全应急响应中心（HUAWEI PSIRT）官方公众号。</p>



  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=7cdf2fb7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FEH8ujwq9pxhT3nQTYokXv6to8cKSnic3oaqY0KmdBQKS3BeCKtr3QudYnWF5ORgEsfgggrJEOC0ic7RAUjkfP0Z8iacwAudrVUiaHtJLO9xymDI%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</p>
  <div style="box-sizing: border-box;font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);" data-pm-slice="0 0 []"><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">引言</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span data-pm-slice="0 0 []"><span leaf="">在《AI时代下对安全攻防演进的思考（上篇）》中，我们看到了AI正成为安全攻防链条中不可忽视的工程力量。本篇将聚焦于此轮变革的核心地带——Web安全、Pwn与逆向工程，探讨AI如何重排这些领域的任务成本与能力门槛。</span></span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">Web 安全：从基础漏洞工业化到 AI 原生应用风险</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">在三大方向中，Web 安全最可能率先发生结构性变化。原因并不复杂：它的输入输出边界相对清晰，反馈速度快，工具链成熟，且大量常见问题具有较强模式性。参数注入是否成功、鉴权是否可绕、回显是否异常、服务端是否触发外联，很多时候只需数轮请求就能得到明确反馈。对人类研究员而言，这意味着密集的重复劳动；对具备脚本生成、上下文记忆与工具调度能力的模型而言，这更像一个适合并行展开的自动循环。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">已发生的事实首先来自厂商一手披露。2025年6月。基于AI驱动的自主渗透测试工具 XBOW 成功登上全球知名漏洞赏金平台 HackerOne 美国排行榜榜首，这是AI系统首次在主流漏洞披露平台上超越人类，标志着网络安全领域的一大里程碑。这些说法不能被直接等同于行业共识，但至少说明：面向真实软件的自动化漏洞发现，已经不再停留在“帮忙写 PoC”或“辅助扫接口”的层面。对 Web 安全而言，这一点尤其敏感，因为浏览器、Web 引擎、网络服务与开放接口本来就是最适合反复探测和验证的对象。</span></p><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">比赛与评测则提供了另一个观察窗口。腾讯云智能渗透挑战赛的公开材料显示，多数参赛方案都把大模型放进多智能体渗透链条之中，用于资产理解、路径探测、工具调度和结果验证。相关报道还提到，基于</span><span lang="EN-US"><span leaf=""> XBOW Benchmark </span></span><span leaf="">的</span><span lang="EN-US"><span leaf=""> 104 </span></span><span leaf="">个漏洞环境中，</span><span lang="EN-US"><span leaf="">XSS</span></span><span leaf="">、默认密码和越权问题占比较高；有队伍将多智能体方案的成功率从</span><span lang="EN-US"><span leaf=""> 50% </span></span><span leaf="">提升到</span><span lang="EN-US"><span leaf=""> 58.2%</span><sup><span leaf="">[8]</span></sup></span><span leaf="">。这些数据当然不能直接代表真实企业环境，却足以说明一个趋势：<span textstyle="" style="font-weight: bold;">越是模式清晰、反馈明确、验证成本低的基础</span></span><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;"> Web </span></span></span><span leaf=""><span textstyle="" style="font-weight: bold;">问题，越容易被</span></span><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;"> AI </span></span></span><span leaf=""><span textstyle="" style="font-weight: bold;">批量吞掉。</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">由此带来的第一个变化，是基础 Web 漏洞会更快暴露，也更快失去稀缺性。过去很多常见问题之所以长期留在系统里，并不完全因为它们隐藏得深，而是因为人工测试产能有限、资产面太散、排查成本过高。AI 介入后，最先贬值的正是这类“低语义、强模式”的漏洞，包括参数污染、简单注入、低阶越权、默认配置、弱口令，以及回显明确的 SSRF 和 XSS 变体。对攻击方而言，发现和利用这类问题的成本还会继续下降；对企业而言，侥幸缓冲期只会越来越短。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">不过，Web 安全并不会因此变得简单，难点只会继续上移。更值钱的部分，将越来越集中在复杂业务逻辑、跨身份边界、多步骤状态切换与信任链组合问题上。优惠券能否叠加、审批流如何回滚、租户边界能否跨越、供应链回调是否可被伪造，这些问题不只是“有没有洞”，更是“系统到底怎么运转、如何授权、责任如何落定”。AI 可以非常高效地撞开很多门，但未必知道哪一扇门背后才是业务的核心利益。但是随着LLM能力的提升，这类业务逻辑漏洞也会慢慢地被自动找出来的，相信这是早晚的事，还有就是业务知识文档的补充，也能在LLM知识不足的情况下，提升LLM对业务逻辑的理解深度，从而进一步挖掘深层次的业务逻辑漏洞。</span></p><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">当前还存在更大的变化，</span><span lang="EN-US"><span leaf="">Web </span></span><span leaf="">应用本身正在迅速演化为</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">应用。开放式</span><span lang="EN-US"><span leaf=""> Web </span></span><span leaf="">应用安全项目（</span><span lang="EN-US"><span leaf="">Open Worldwide Application Security Project</span></span><span leaf="">，</span><span lang="EN-US"><span leaf="">OWASP</span></span><span leaf="">）的</span><span lang="EN-US"><span leaf=""> GenAI Security Project </span></span><span leaf="">已将</span><span lang="EN-US"><span leaf=""> Agentic App Security</span></span><span leaf="">、</span><span lang="EN-US"><span leaf="">Red Teaming &amp; Evaluation</span></span><span leaf="">、数据安全与治理列为独立方向；美国国家标准与技术研究院（</span><span lang="EN-US"><span leaf="">National Institute of Standards and Technology</span></span><span leaf="">，</span><span lang="EN-US"><span leaf="">NIST</span></span><span leaf="">）的《</span><span lang="EN-US"><span leaf="">Generative AI Profile</span></span><span leaf="">》也强调，生成式</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">的风险管理必须贯穿设计、开发、使用与评估全过程</span><sup><span lang="EN-US"><span leaf="">[12][13]</span></span></sup><span leaf="">。落到</span><span lang="EN-US"><span leaf=""> Web </span></span><span leaf="">体系中，这意味着安全对象正在从</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">请求、响应、数据库</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">扩展到</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">提示词、上下文、检索数据、工具权限以及模型驱动的动作链</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。过去重点防守的是</span><span lang="EN-US"><span leaf=""> XSS</span></span><span leaf="">、</span><span lang="EN-US"><span leaf="">CSRF</span></span><span leaf="">、</span><span lang="EN-US"><span leaf="">SSRF</span></span><span leaf="">、越权；接下来还必须同时面对</span><span lang="EN-US"><span leaf=""> prompt injection</span></span><span leaf="">、</span><span lang="EN-US"><span leaf="">tool abuse</span></span><span leaf="">、上下文污染、跨插件信任边界失守与代理越权执行。也正因此，未来两三年</span><span lang="EN-US"><span leaf="">Web </span></span><span leaf="">安全最明显的变化，未必只是技术手段更新，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">发现漏洞、理解漏洞、处理漏洞</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">的整套节奏被重新定义。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><strong style="box-sizing: border-box;"><span leaf="">Pwn：前段自动化与后段高手化并存</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">Pwn 领域真正发生变化的，不是“模型会不会一把打穿系统”，而是漏洞研究流程正在被重新切段。就目前公开材料看，AI 最先接管的不是终局利用，而是前面的漏洞挖掘、variant analysis、崩溃归因、PoC 生成和补丁回归检查；难的部分，仍然是把一个 bug 稳定、低成本、跨环境地变成任意的控制能力。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">从公开披露来看，</span><span lang="EN-US"><span leaf="">Google Project Naptime </span></span><span leaf="">在</span><span lang="EN-US"><span leaf=""> 2024 </span></span><span leaf="">年已经明确展示出一个很有现实意义的方向：围绕已知补丁、崩溃和代码上下文持续做</span><span lang="EN-US"><span leaf=""> variant analysis</span><sup><span leaf="">[1]</span></sup></span><span leaf="">。这意味着</span><span lang="EN-US"><span leaf=""> Pwn </span></span><span leaf="">中最耗时间的前置劳动正在被压缩。</span><span lang="EN-US"><span leaf="">Big Sleep </span></span><span leaf="">进一步把这一方向推到真实代码库。</span><span lang="EN-US"><span leaf="">2024 </span></span><span leaf="">年</span><span lang="EN-US"><span leaf=""> 11 </span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Project Zero </span></span><span leaf="">披露其在</span><span lang="EN-US"><span leaf=""> SQLite </span></span><span leaf="">中发现了此前未知、可利用的内存安全漏洞，并在发布前完成报告与修复；到</span><span lang="EN-US"><span leaf="">2025 </span></span><span leaf="">年</span><span lang="EN-US"><span leaf=""> 7 </span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Google </span></span><span leaf="">又称相关系统已能结合威胁情报，提前拦截即将进入在野利用阶段的漏洞</span><sup><span lang="EN-US"><span leaf="">[2][3]</span></span></sup><span leaf="">。这些说法仍属于厂商一手材料，外部难以完全复核，但至少揭示出一个越来越清楚的趋势：<span textstyle="" style="font-weight: bold;">从补丁线索到可利用变种的时间窗正在缩短。</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span lang="EN-US"><span leaf="">Anthropic </span></span><span leaf="">的披露把能力边界说得更具体。根据其</span><span lang="EN-US"><span leaf=""> 2026</span></span><span leaf="">年</span><span lang="EN-US"><span leaf="">3</span></span><span leaf="">月与</span><span lang="EN-US"><span leaf=""> Mozilla </span></span><span leaf="">的公开合作材料，</span><span lang="EN-US"><span leaf="">Claude </span></span><span leaf="">在大规模</span><span lang="EN-US"><span leaf=""> C++ </span></span><span leaf="">代码中提交了</span><span lang="EN-US"><span leaf=""> 112 </span></span><span leaf="">份唯一报告，其中</span><span lang="EN-US"><span leaf="">22</span></span><span leaf="">个被认定为漏洞，</span><span lang="EN-US"><span leaf="">14</span></span><span leaf="">个为高危；但在利用侧，只有极少数尝试被转成可运行</span><span lang="EN-US"><span leaf=""> exploit</span></span><span leaf="">，而且成立条件是刻意削弱防护的测试环境</span><sup><span lang="EN-US"><span leaf="">[14]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">Claude Mythos Preview </span></span><span leaf="">进一步展示了模型在</span><span lang="EN-US"><span leaf=""> Linux </span></span><span leaf="">内核</span><span lang="EN-US"><span leaf=""> N-day </span></span><span leaf="">筛选、漏洞串联和本地提权上的跃迁，同时也承认远程触发、真实缓解机制和复杂环境适配仍是明显短板</span><sup><span lang="EN-US"><span leaf="">[5]</span></span></sup><span leaf="">。将这几份材料合在一起看，</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">在</span><span lang="EN-US"><span leaf=""> Pwn </span></span><span leaf="">中更擅长的是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">找、读、试、归纳</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，弱项仍然是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">稳、深、通杀</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">因此，Pwn 的未来更像“前半程工业化，后半程继续高手化”。漏洞挖掘、变体分析、PoC 生成、patch validation，非常适合交给 AI 做高通量筛查；缓解绕过、堆布局塑形、竞态利用、跨环境适配和链式提权，仍然高度依赖研究员对体系结构、编译器与运行时语义的把握。对企业而言，这意味着补丁响应、回归验证、sanitizer 部署和 memory-safe 迁移会比以往更重要；对研究员而言，稀缺能力会进一步上移到利用面判断、验证路径设计和复杂 exploitation 上。从Anthropic 博客上提到的 Firefox 漏洞利用事件上看，AI也正在尝试蚕食pwn的后半程研究，只不过当前还没达到通用稳定、全自动、普及化的程度，研究员在后半程高手化的持续时间也许不会太久。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><strong style="box-sizing: border-box;"><span leaf="">逆向工程：分析自动化与抗分析保护的相互抬升</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">如果说 Web 安全最先感受到的是“基础漏洞被批量吞掉”，那么逆向工程感受到的，更多是一种分析方式本身的改写。过去做逆向，最耗时间的往往不是最后那一下定性判断，而是前面那一长串脏活：整理反汇编结果、补函数语义、切换 Ghidra、IDA、调试器和脚本、追踪可疑分支、抽取配置、猜测协议字段。过去两年里，一个可以确认的事实是，大模型已经开始显著压缩这部分前置工作。它不只是把汇编翻成更像 C 的伪代码，更开始围绕分析目标组织工具、生成脚本、归纳局部语义，再把这些碎片拼成一份初步可读的理解框架。</span></p><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">这方面最有代表性的工作之一，2024</span><span leaf="">年</span><span lang="EN-US"><span leaf="">EMNLP</span></span><span leaf="">的《</span><span lang="EN-US"><span leaf="">LLM4Decompile</span></span><span leaf="">》。这项研究讨论的不是</span><span lang="EN-US"><span leaf="">“AI </span></span><span leaf="">会不会写代码</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是二进制反编译能否从传统工具生成的低可读、低可执行伪代码，迈向更接近源码层面的恢复结果。论文给出的结论是，在其设定的</span><span lang="EN-US"><span leaf=""> HumanEval </span></span><span leaf="">和</span><span lang="EN-US"><span leaf=""> ExeBench </span></span><span leaf="">基准上，专门训练的模型在可执行性和可读性上明显优于</span><span lang="EN-US"><span leaf=""> Ghidra </span></span><span leaf="">和通用模型</span><sup><span lang="EN-US"><span leaf="">[15]</span></span></sup><span leaf="">。这当然不能直接外推成</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">现实世界中的复杂二进制已经可以被稳定自动还原</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，因为论文环境仍是受控评测；但它至少说明，逆向工程中长期被视作纯手工活的</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">语法恢复</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">和</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">局部语义修补</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，已经开始出现可重复的模型收益。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">比单点反编译能力更值得注意的，是围绕逆向任务的工具编排能力。模型一旦接上反汇编器、调试器、字符串提取器、沙箱和脚本执行环境，它做的就不再只是“翻译”汇编，而会进入一个更像分析员的循环：定位入口和关键 API，决定先静态看调用链还是先跑起来抓行为，遇到字符串加密就先解一层，遇到配置块就尝试提取结构。恶意代码分析因此最先受益。对中低复杂度样本而言，AI 完全可能在较短时间内完成家族初筛、配置提取、通信逻辑粗分和可疑函数定位，把原本需要资深分析员花数小时完成的整理工作压缩掉大半。</span></p><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">但逆向工程与</span><span lang="EN-US"><span leaf=""> Web </span></span><span leaf="">测试、代码审计有一个根本差别：它面对的对象往往不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">自然形成的复杂性</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">刻意制造的不可读性</span><span lang="EN-US"><span leaf="">”。</span></span><span lang="EN-US"><span leaf="">2025</span></span><span leaf="">年的《</span><span lang="EN-US"><span leaf="">Deconstructing Obfuscation</span></span><span leaf="">》给出了一个重要边界。作者评估了多种商业模型在汇编去混淆任务中的表现，发现模型面对低强度噪声并非完全无能，但一旦进入控制流平坦化、指令替换，尤其是多种混淆技术叠加的场景，性能会明显下滑。论文总结出的失败模式也很典型，包括谓词误判、控制流映射错误、算术变换理解偏差和常量传播失真</span><sup><span lang="EN-US"><span leaf="">[16]</span></span></sup><span leaf="">。换句话说，<span textstyle="" style="font-weight: bold;">模型并不是</span></span><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;">“</span></span></span><span leaf=""><span textstyle="" style="font-weight: bold;">看不懂汇编</span></span><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;">”</span></span></span><span leaf=""><span textstyle="" style="font-weight: bold;">，它更常见的问题是：会给出一套表面自洽、局部合理、但在整体执行语义上站不住的解释。</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">这正是为什么虚拟机保护、控制流平坦化、动态解密、反调试和环境绑定在 AI 时代反而会变得更重要。它们的价值，从来不只是把人工分析拖慢，而是切断语义连续性，让分析者无法轻易把局部模式拼成完整逻辑。对 LLM 来说，这类保护尤为致命，因为<span textstyle="" style="font-weight: bold;">模型的优势往往建立在“模式可识别、上下文可延展、局部语义能稳定累积”这几个前提上。一旦执行路径被拆碎、关键常量运行时生成、核心逻辑被藏进自定义虚拟机，模型就很容易陷入一种危险状态：它仍能生成流畅解释，但解释越来越像伪语义，而不是程序真正执行的语义。</span></span></p><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span lang="EN-US"><span leaf="">NDSS BAR 2026 </span></span><span leaf="">的《</span><span lang="EN-US"><span leaf="">Towards LLM-Resistant Software Protection</span></span><span leaf="">》又将这一问题往前推了一步。严格地说，这类研究更多是在</span><span lang="EN-US"><span leaf=""> CTF </span></span><span leaf="">逆向题环境中观察</span><span lang="EN-US"><span leaf=""> agent </span></span><span leaf="">的失败模式，还不能直接当作产业结论；但其释放出的信号已经非常清晰：<span textstyle="" style="font-weight: bold;">未来的软件保护，很可能会开始有意识地针对</span></span><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;"> AI agent </span></span></span><span leaf=""><span textstyle="" style="font-weight: bold;">的弱点进行设计，而不只是针对人类分析员</span></span><sup><span lang="EN-US"><span leaf=""><span textstyle="" style="font-weight: bold;">[17]</span></span></span></sup><span leaf=""><span textstyle="" style="font-weight: bold;">。</span>如果把这些事实和论文结论放在一起看，一个相对稳妥的趋势判断是：逆向工程不会因为</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">变强而迅速失去专业性，它更可能出现明显分化。一端是入门和中段分析被显著加速，样本分流、配置提取、基础反编译、常规恶意代码初筛越来越自动化；另一端则是保护工程、去混淆、动态行为复原和最终定性判断变得更加值钱。逆向工程因此很可能成为</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">时代最早出现新军备竞赛特征的安全子领域之一。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><strong style="box-sizing: border-box;"><span leaf="">跨领域迁移：漏洞类型、防守重心与组织能力重排</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">将 Web 安全、Pwn 与逆向工程放在一起观察，可以看到一个相当一致的结构性变化： AI 最先吞掉的，往往不是最复杂的问题，而是那些可以被表达、可以被验证、可以被批量搜索的问题。模式性强、反馈短、工具成熟的任务，会优先进入自动化扩张区；需要复杂上下文、长期试错、强责任判断和高对抗性的任务，则会进一步成为稀缺能力。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">这意味着基础漏洞不会消失，但会更早暴露、更快被利用，也会更快失去稀缺性。<span textstyle="" style="font-weight: bold;">对于攻击者而言，进入门槛在下降；对于防守者而言，侥幸拖延的空间在收缩。相应地，高价值问题会向更复杂的层次迁移：Web 领域向业务逻辑、身份边界和 AI 原生应用调用链迁移，Pwn 领域向稳定利用、缓解绕过与环境适配迁移，逆向领域向混淆保护、动态行为理解与抗分析工程迁移。</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">还要看到，组织能力的重心也在改变。过去安全团队常把大量资源投入在“找到更多问题”上；未来更关键的问题可能是，如何在 AI 不断扩大问题发现规模的同时，维持足够快的 triage、复现、修补、回归、上线和审计节奏。如果没有后续流程支撑，再强的发现能力最终也只会演变为告警洪水。也正因此，<span textstyle="" style="font-weight: bold;">AI 对安全攻防的真正冲击，并不是工具层面的单点替代，而是迫使企业重新设计验证、授权、补丁、日志和责任链条。</span></span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">小结</span></strong></p></div></div></div></div></div><span style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">      AI正在将安全攻防推向“基础问题工业化，复杂问题上移”的新阶段。当发现问题的能力被极大增强后，真正的挑战随之而来：企业、安全人员乃至整个行业，应如何应对这场系统性变革？未来的竞争格局将如何演变？《AI时代下对安全攻防演进的思考（下篇）》将继续分析。</span></span><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">参考文献</span></strong></p></div></div></div></div></div><div style="text-align: center;box-sizing: border-box;"><div style="display: inline-block;width: 100%;height: 240px;vertical-align: top;overflow-y: auto;box-sizing: border-box;"><p style="font-size: 12px;text-align: left;box-sizing: border-box;"><ol style="list-style-type: decimal;box-sizing: border-box;padding-left: 20px;list-style-position: outside;" class="list-paddingleft-2"><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，Project Naptime: Evaluating      Offensive Security Capabilities of Large Language Models，2024-06-20。<a href="https://projectzero.google/2024/06/project-naptime.html" target="_blank">https://projectzero.google/2024/06/project-naptime.html</a></span></p></li><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，From Naptime to Big Sleep: Using      Large Language Models To Catch Vulnerabilities In Real-World Code，2024-11-01。<a href="https://projectzero.google/2024/10/from-naptime-to-big-sleep.html" target="_blank">https://projectzero.google/2024/10/from-naptime-to-big-sleep.html</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google，A summer of security: empowering cyber defenders      with AI，2025-07-15。<a href="https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/" target="_blank">https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Project Glasswing: Securing critical software for      the AI era，2026-04-07。<a href="https://www.anthropic.com/glasswing" target="_blank">https://www.anthropic.com/glasswing</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude Mythos Preview，2026-04-07。<a href="https://red.anthropic.com/2026/mythos-preview/" target="_blank">https://red.anthropic.com/2026/mythos-preview/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Occupational changes during the      20th century，2006-03。<a href="https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm" target="_blank">https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Federal Highway Administration，The Evolution of MUTCD。<a href="https://mutcd.fhwa.dot.gov/kno-history.htm" target="_blank">https://mutcd.fhwa.dot.gov/kno-history.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">腾讯云黑客松官网；腾讯云开发者社区《400+极客菁英共聚羊城，见证国内首个AI智能渗透挑战赛》。<a href="https://tch.cloud.tencent.com/" target="_blank">https://tch.cloud.tencent.com/</a> ；<a href="https://cloud.tencent.com/developer/article/2651925" target="_blank">https://cloud.tencent.com/developer/article/2651925</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，AIxCC:      AI Cyber Challenge；AI Cyber Challenge marks      pivotal inflection point for cyber defense。<a href="https://www.darpa.mil/research/programs/ai-cyber" target="_blank">https://www.darpa.mil/research/programs/ai-cyber</a> ；<a href="https://www.darpa.mil/news/2025/aixcc-results" target="_blank">https://www.darpa.mil/news/2025/aixcc-results</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，Cyber      Grand Challenge (CGC)。<a href="https://www.darpa.mil/research/programs/cyber-grand-challenge" target="_blank">https://www.darpa.mil/research/programs/cyber-grand-challenge</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude is competitive with humans      in (some) cyber competitions，2025-08-09。<a href="https://red.anthropic.com/2025/cyber-competitions/" target="_blank">https://red.anthropic.com/2025/cyber-competitions/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">National Institute of Standards and      Technology (NIST)，Artificial      Intelligence Risk Management Framework: Generative Artificial Intelligence      Profile，2024-07-26。<a href="https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence" target="_blank">https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Open Worldwide Application Security      Project (OWASP) GenAI Security Project，OWASP Top 10 for LLM is now the GenAI Security      Project and promoted to OWASP Flagship status，2025-03-26。<a href="https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/" target="_blank">https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Partnering with Mozilla to      improve Firefox&#39;s security，2026-03-06。<a href="https://red.anthropic.com/2026/firefox/" target="_blank">https://red.anthropic.com/2026/firefox/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Hanzhuo Tan, Qi Luo, Jing Li, Yuqun      Zhang，LLM4Decompile:      Decompiling Binary Code with Large Language Models，EMNLP 2024。<a href="https://aclanthology.org/2024.emnlp-main.203/" target="_blank">https://aclanthology.org/2024.emnlp-main.203/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anton Tkachenko, Dmitrij Suskevic,      Benjamin Adolphi，Deconstructing      Obfuscation: A four-dimensional framework for evaluating Large Language      Models assembly code deobfuscation capabilities，2025。<a href="https://arxiv.org/abs/2505.19887" target="_blank">https://arxiv.org/abs/2505.19887</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Abraham Clements et al.，Towards LLM-Resistant Software      Protection: Agent Failure Patterns in CTF Reverse Engineering，NDSS BAR 2026。<a href="https://www.ndss-symposium.org/ndss-paper/auto-draft-657/" target="_blank">https://www.ndss-symposium.org/ndss-paper/auto-draft-657/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Cybench: A Framework for Evaluating      Cybersecurity Capabilities and Risks of Language Models。<a href="https://cybench.github.io/" target="_blank">https://cybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">BountyBench: Dollar Impact of AI Agent      Attackers and Defenders on Real-World Cybersecurity Systems。<a href="https://bountybench.github.io/" target="_blank">https://bountybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Coordinated vulnerability disclosure for      Claude-discovered vulnerabilities，最后更新于 2026-03-06。<a href="https://www.anthropic.com/coordinated-vulnerability-disclosure" target="_blank">https://www.anthropic.com/coordinated-vulnerability-disclosure</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Node.js Project，New HackerOne Signal Requirement      for Vulnerability Reports，最后更新于 2026-02-19。<a href="https://nodejs.org/en/blog/announcements/hackerone-signal-requirement" target="_blank">https://nodejs.org/en/blog/announcements/hackerone-signal-requirement</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">HackerOne，Code of Conduct。<a href="https://www.hackerone.com/policies/code-of-conduct" target="_blank">https://www.hackerone.com/policies/code-of-conduct</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Daniel Stenberg，The end of the curl bug-bounty，2026-01-26。<a href="https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/" target="_blank">https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">前沿观察 赛事纪实：从腾讯云黑客松，洞见智能体时代的攻防新格局，2026-04-17。<a href="https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q" target="_blank">https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Changes in the U.S. occupational      mix from 1860 to 2015，2019-08。<a href="https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm" target="_blank">https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Martin Fiszbein, Jeanne Lafortune, Ethan      G. Lewis, José Tessada，Powering Up Productivity: The Effects of Electrification on      U.S. Manufacturing，National Bureau of Economic      Research (NBER) Working Paper 28076，2020；2024-04 修订。<a href="https://www.nber.org/papers/w28076" target="_blank">https://www.nber.org/papers/w28076</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">James Bessen，Toil and Technology，International Monetary Fund (IMF) Finance &amp; Development，2015-03。<a href="https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm" target="_blank">https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Timothy F. Bresnahan, Manuel Trajtenberg，General Purpose Technologies      &#34;Engines of Growth?&#34;，National Bureau of      Economic Research (NBER) Working Paper 4148，1992-08。<a href="https://www.nber.org/papers/w4148" target="_blank">https://www.nber.org/papers/w4148</a></span></p></li></ol></p></div></div></div><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box;overflow-wrap: break-word !important;clear: both;min-height: 1em;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;" data-pm-slice="6 2 []"><strong data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;letter-spacing: 0.544px;display: inline;color: rgb(62, 62, 62);font-family: 楷体;text-align: left;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;letter-spacing: normal;color: rgb(255, 0, 0);font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><strong style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;display: inline;color: rgb(255, 79, 121);font-size: 16px;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><span style="color: rgb(62, 62, 62);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 16px;font-style: italic;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;display: inline !important;float: none;" data-pm-slice="0 0 []"><span leaf="">系列连载未完待续，下篇敬请期待。</span></span></span></strong></span></strong></p><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box;overflow-wrap: break-word !important;clear: both;min-height: 1em;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;" data-pm-slice="6 2 []"><strong data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;letter-spacing: 0.544px;display: inline;color: rgb(62, 62, 62);font-family: 楷体;text-align: left;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;letter-spacing: normal;color: rgb(255, 0, 0);font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><strong style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;display: inline;color: rgb(255, 79, 121);font-size: 16px;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">本公众号发布、转载的文章所涉及的技术、思路、工具仅供学习交流，任何人不得将其用于非法用途及盈利等目的，否则后果自行承担！</span></span></strong></span></strong></p><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;text-align: center;"><span data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">点这里 <img data-aistatus="1" alt="图片" class="rich_pages wxw-img __bg_gif" data-ratio="0.5982532751091703" data-w="458" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;vertical-align: middle;height: auto !important;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;font-family: PingFangSC-Regular, &#34;PingFang SC&#34;;text-indent: 28px;color: rgb(62, 62, 62);font-size: 16px;width: 61.9922px !important;visibility: visible !important;" data-width="100%" data-imgfileid="100041710" src="https://wechat2rss.xlab.app/img-proxy/?k=6849f525&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_gif%2FMfTd6rd9CyvNRMW8I9cvI1CK5gKiaYqg2veTn9t9dAe1GxYic7pAvgvRIKNFickConFyX8AvW2reAq8GchJI6aBpA%2F640%3Fwx_fmt%3Dgif%26wxfrom%3D5%26wx_lazy%3D1%26tp%3Dwebp%23imgIndex%3D14"/></span><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" data-pm-slice="1 1 [&#34;para&#34;,null,&#34;node&#34;,{&#34;tagName&#34;:&#34;span&#34;,&#34;attributes&#34;:{&#34;style&#34;:&#34;color: rgb(127, 229, 230); font-family: \&#34;Helvetica Neue\&#34;, Helvetica, \&#34;Hiragino Sans GB\&#34;, \&#34;Microsoft YaHei\&#34;, Arial, sans-serif; font-size: 14px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: 0.578px; orphans: 2; text-align: center; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px;  background-color: rgb(255, 255, 255); text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; display: inline !important; float: none;&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;}]" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">关注我们，一键三连～</span></span></span></p><p class="mp_profile_iframe_wrp" nodeleaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;"><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe js_wx_tap_highlight" data-pluginname="mpprofile" data-nickname="华为安全应急响应中心" data-alias="HUAWEI_PSIRT" data-index="0" data-from="2" data-headimg="http://mmbiz.qpic.cn/sz_mmbiz_png/Pf9eicDVDMxHbPW1POGs9HHQCUGXBXg7u6TCtI2ab5DdIEfxJWcR46krXgudVuibfibsqRYlAtN2RLdaiaOCosQMSw/300?wx_fmt=png&amp;wxfrom=19" data-signature="华为安全应急响应中心（HUAWEI PSIRT）官方公众号。" data-id="MzI0MTY5NDQyMw==" data-is_biz_ban="0" data-origin_num="49" data-biz_account_status="0" data-service_type="1" data-verify_status="2"></mp-common-profile></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=3799aebe&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486092%26idx%3D2%26sn%3Dc494c0340d1a05ef458bec7584c994bf">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 09 Jul 2026 18:02:00 +0800</pubDate>
    </item>
    <item>
      <title>AI时代下对安全攻防演进的思考（下篇）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486092&amp;idx=3&amp;sn=1fdedc10e9db976f0d01304dd4490da7</link>
      <description>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</description>
      <content:encoded><![CDATA[<p><span>riusksk</span> <span>2026-07-09 18:02</span> <span style="display: inline-block;">广东</span></p>




  <p>以下文章来源于：华为安全应急响应中心</p>
  <strong>华为安全应急响应中心</strong>
  <p>华为安全应急响应中心（HUAWEI PSIRT）官方公众号。</p>



  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=3ca143a7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FEH8ujwq9pxiaxeMmre0w0ouERuUkFUeDHCgDE3RZfB1mK1pEf5Uy3jtPFTGzticlaeXCUb786tZDRVDYkFibJlUoWGK4WDvbQndc3QqN1DB5hY%2F0%3Fwx_fmt%3Djpeg"/></p>
  <p>AI正重构安全攻防成本结构：基础漏洞发现、变体分析等环节已获效率提升，但复杂利用与责任判断仍依赖人力。这推动企业安全建设从“多发现问题”转向“快验证、快修复、严治理”，安全人员需上移至系统理解、流程编排等更高阶能力。攻防重心正向纵深治理迁移。</p>
  <div style="box-sizing: border-box;font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);" data-pm-slice="11 8 []"><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">引言</span></strong></p></div></div></div></div></div><span style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);box-sizing: border-box;text-indent: 2em;" data-pm-slice="0 0 []"><span leaf="">     上篇与中篇分别探讨了AI赋能安全攻防的现状及其在三大核心领域引发的结构性变化。本篇作为系列终章，将深入分析这一变革对企业运营、个人发展及行业生态带来的深远影响，并展望未来的攻防重心。</span></span><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">影响分析</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p data-pm-slice="0 0 []" style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><b><span leaf="">1、对企业意味着什么：比工具升级更难的是流程重写</span></b></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">对企业而言，最直接的变化并不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">是否采购大模型</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">现有安全流程能否承受更高频、更大规模的问题发现能力</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。如果说过去很多企业面对安全工具时，主要解决的是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">要不要多上一套扫描、审计或渗透能力</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，那么在</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">进入攻防链条之后，难的部分已经不再是工具采购，而是流程重写。模型可以更快发现问题，却不会自动替企业完成定级、复现、修复、回归、审批和审计。发现能力一旦扩张，而后续流程没有同步升级，企业得到的往往不是更强的防御，而是更密集的告警、更多相互竞争的优先级，以及更容易失控的修复节奏。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">因此，企业首先被迫改变的，将是漏洞响应机制。高危问题的确认、责任归属、修复窗口、变更验证和上线回归，都需要比过去更快地闭环。第二个必须前置的能力，是验证能力。未来企业安全团队的重要职责，不再只是亲手把问题找出来，而是能够迅速判断</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">报告是否正确、影响范围有多大、应当由谁处理、是否需要阻断发布或扩大排查。第三个变化是授权与审计边界必须更细。只要</span><span lang="EN-US"><span leaf=""> Agent </span></span><span leaf="">开始参与测试、代码修改、补丁生成甚至自动执行任务，工具权限、日志留痕、审批节点和回滚机制就不能继续停留在粗放状态。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">这类约束在大型企业中具有普遍性。对于组织规模大、研发链路长、信息安全要求高的企业而言，外部顶尖闭源模型并不总能直接接入真实研发与安全流程；即便企业选择自建或微调开源模型，也往往还要同时解决私域知识接入、权限控制、推理成本、公司级算力供给、内网隔离和跨团队协作接口改造等问题。大企业船大难调头，难的通常不是单个模型效果，而是整套研发与安全流程是否愿意、也是否有能力围绕新工具重新组织。企业真正需要准备的不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">接入</span><span lang="EN-US"><span leaf=""> AI”</span></span><span leaf="">，而是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">接住</span><span lang="EN-US"><span leaf=""> AI”</span></span><span leaf="">。流程是否重写，将直接决定</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">对企业来说究竟是防御增益，还是新的管理负担。如果后续</span><span lang="EN-US"><span leaf=""> triage</span></span><span leaf="">、修复、回归和审计跟不上，再强的问题发现能力最终都可能变成另一种形式的噪音生产线。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf=""><span textstyle="" style="font-weight: bold;">2、对安全人员的影响</span></span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">对安全人员而言，中低端、重复型、模式化工作被压缩几乎已成定局。单纯依赖熟练使用少量工具、背诵常见</span><span lang="EN-US"><span leaf=""> payload</span></span><span leaf="">、重复执行既定流程所形成的比较优势，会变得越来越薄。与此同时，人的高价值能力并没有消失，反而更集中在四个方面：</span></p><ul style="font-style: normal;font-weight: 400;text-align: justify;font-size: 16px;color: rgb(62, 62, 62);white-space: normal;box-sizing: border-box;text-indent: 2em;" class="list-paddingleft-1"><li><p style="text-indent: 0px;"><span leaf="">理解复杂系统与真实业务约束；</span></p></li><li><p style="text-indent: 0px;"><span leaf="">组织上下文、多种工具协同、工作流编排；</span></p></li><li><p style="text-indent: 0px;"><span leaf="">验证与反驳</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">生成的结果；</span></p></li><li><p style="text-indent: 0px;"><span leaf="">在披露、修补和治理中承担责任。</span></p></li></ul><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">从近年的实战与竞赛反馈看，</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">对安全行业带来的并不是简单的</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">平权</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">或</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">替代</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而更像是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">低门槛、高天花板</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。一方面，入门门槛确实下降了。零基础或低经验参与者可以在</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">辅助下更快搭建</span><span lang="EN-US"><span leaf=""> agent</span></span><span leaf="">、理解基本攻击链条、完成以赛促学的第一轮成长。另一方面，真正的差距并没有消失，而是被重新定义并进一步放大。谁能更准确地拆解任务、设计</span><span lang="EN-US"><span leaf=""> agent </span></span><span leaf="">架构、选择工具组合、控制</span><span lang="EN-US"><span leaf=""> token </span></span><span leaf="">成本、处理错误恢复和长链路调度，谁就更容易把</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">变成稳定生产力。这也是为什么安全人员的角色正在从</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">执行者</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">转向</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">指挥者</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">：未来更有价值的，不是亲手完成每一步重复操作的人，而是能够定义问题、拆解目标、调度</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">与工具并对最终结果负责的人。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">如果把这一变化再往深处看，会发现模型能力本身始终有上限，而知识没有上限。模型可以随着参数、工具和训练不断增强，但它在某个时刻能够稳定处理的问题范围、上下文容量和推理质量，仍然受制于架构、成本和工程条件；相比之下，领域知识、实战经验、案例积累、系统理解和技术审美，并不存在同样明确的天花板。也正因为如此，</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">时代对人类真正稀缺的不是单纯</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">谁先拿到一个更强模型</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，还考验谁拥有更深的领域知识，谁能把长期沉淀的专家经验转化成更好的问题定义、更合理的约束设计和更可靠的验证路径。模型是杠杆，但知识深度才是支点。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">更进一步说，问题的定义能力与结果的验证能力，正在从安全岗位的专业要求，变成</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">时代每个人都应具备的基础能力。近一年的公开事件已经说明，如果使用者不能清晰定义问题、约束模型输出并亲自验证结果，</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">往往不会把高质量能力普及给更多人，反而会把低质量内容更快地批量制造出来。</span><span lang="EN-US"><span leaf="">2026 </span></span><span leaf="">年</span><span lang="EN-US"><span leaf=""> 2 </span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">Node.js </span></span><span leaf="">官方宣布将</span><span lang="EN-US"><span leaf="">HackerOne </span></span><span leaf="">提交门槛提高到</span><span lang="EN-US"><span leaf=""> Signal 1.0</span></span><span leaf="">（高质量漏洞报告），理由是安全团队经历了因使用</span><span lang="EN-US"><span leaf="">AI</span></span><span leaf="">生成导致显著增加的低质量报告</span><sup><span lang="EN-US"><span leaf="">[21]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">HackerOne </span></span><span leaf="">在官方行为准则中进一步明确要求</span><span lang="EN-US"><span leaf=""> human-in-the-loop</span></span><span leaf="">，强调研究者必须对</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">生成内容负责，并禁止大规模提交</span><span lang="EN-US"><span leaf=""> low-signal</span></span><span leaf="">、未经验证或缺乏</span><span lang="EN-US"><span leaf=""> PoC </span></span><span leaf="">的报告</span><sup><span lang="EN-US"><span leaf="">[22]</span></span></sup><span leaf="">。</span><span lang="EN-US"><span leaf="">2026 </span></span><span leaf="">年</span><span lang="EN-US"><span leaf=""> 1 </span></span><span leaf="">月，</span><span lang="EN-US"><span leaf="">curl </span></span><span leaf="">项目则更进一步，直接结束</span><span lang="EN-US"><span leaf=""> bug bounty</span></span><span leaf="">，并公开将</span><span lang="EN-US"><span leaf="">“AI slop reports”</span></span><span leaf="">列为关键原因之一，同时指出其</span><span lang="EN-US"><span leaf=""> 2025 </span></span><span leaf="">年确认有效漏洞的比例已跌破</span><span lang="EN-US"><span leaf=""> 5%</span><sup><span leaf="">[23]</span></sup></span><span leaf="">。这些事件背后暴露的不是单纯的社区管理问题，而是一个更普遍的现实：</span><b><span leaf="">如果不会定义问题、不会验证结果，</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">只会放大噪音，而不会稳定放大能力。</span></b></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">因此，未来更值钱的安全人员，不是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">最会操作工具</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">的人，而是能够在</span><span lang="EN-US"><span leaf=""> AI</span></span><span leaf="">、代码、业务、架构与风险之间做出可靠判断的人。模型会普及一些过去需要经验积累才能获得的能力，但也会把真正难的部分抬得更高。问题定义与结果验证，正是其中最基础、也最不可外包给模型的两项能力。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><b><span leaf="">3、对行业发展的影响</span></b></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">对行业而言，三到五年内最显著的变化可能有三点。第一，围绕</span><span lang="EN-US"><span leaf=""> Web </span></span><span leaf="">渗透、代码审计、</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">应用安全评估的产品化竞争会明显加速，基础服务将被重新定价。第二，</span><span lang="EN-US"><span leaf="">Pwn </span></span><span leaf="">方向会继续沿</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">自动发现加速、自动利用谨慎推进</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">的路径发展，自动</span><span lang="EN-US"><span leaf=""> triage</span></span><span leaf="">、自动</span><span lang="EN-US"><span leaf="">patching </span></span><span leaf="">和</span><span lang="EN-US"><span leaf=""> memory-safe </span></span><span leaf="">迁移支持会变得更重要。第三，逆向工程将从效率提升进一步走向</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">混淆保护对抗</span><span lang="EN-US"><span leaf="">AI</span></span><span leaf="">自动化分析</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">的新型竞争。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">与此同时，工具本身的价值结构也会发生变化。传统意义上</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">为人直接操作</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">而设计的大量胶水型脚本和弱封装工具，会逐步被更高层的</span><span lang="EN-US"><span leaf=""> agent </span></span><span leaf="">编排吸收；真正长期值钱的，将是那些能够提供确定性执行、稳定发包、标准接口、底层协议处理和可靠基础能力的工具组件。换句话说，工具不再只是</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">给人用</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，而会越来越多地变成</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">给</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">调用</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。这也解释了为什么原子化能力、标准化接口、面向函数调用和</span><span lang="EN-US"><span leaf=""> MCP </span></span><span leaf="">等协议的设计，以及工程化兜底机制，会在</span><span lang="EN-US"><span leaf=""> AI </span></span><span leaf="">时代变得比过去更重要。</span><span lang="EN-US"><span leaf="">AI </span></span><span leaf="">负责不确定性探索，工具负责确定性执行，二者的边界越清晰，系统整体越可控。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">因此，行业竞争最终不会只体现在</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">谁的模型更强</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">，还会体现在</span><span lang="EN-US"><span leaf="">“</span></span><span leaf="">谁的工具生态更适合被模型调用</span><span lang="EN-US"><span leaf="">”“</span></span><span leaf="">谁能把幻觉和不稳定性压到工程可接受范围内</span><span lang="EN-US"><span leaf="">”“</span></span><span leaf="">谁能形成从模型到接口、从接口到流程的标准化闭环</span><span lang="EN-US"><span leaf="">”</span></span><span leaf="">。</span><span lang="EN-US"><span leaf="">Cybench</span></span><span leaf="">、</span><span lang="EN-US"><span leaf="">BountyBench</span></span><span leaf="">这类评测框架的重要性也会持续提升，因为行业越来越需要把厂商叙事、竞赛成绩与现实部署能力区分开来。没有统一或至少可对照的评测体系，企业很难判断哪些是演示，哪些是能力；哪些是一次性样板，哪些是真正可以稳定纳入流程的工程对象</span><sup><span lang="EN-US"><span leaf="">[18] [19]</span></span></sup><span leaf="">。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">结论</span></strong></p></div></div></div></div></div><div style="box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">放到今天看，AI 对安全攻防的影响已经越过“辅助工具”阶段，但还远没到“全面替代”阶段。更接近事实的说法是，AI 正在重排安全任务的成本结构，也在重排漏洞发现、漏洞利用、逆向分析和安全治理中的稀缺能力分布。Web 安全会更快走向基础问题工业化，Pwn 会更明显地分化出“前段自动化、后段高门槛”，逆向工程则会进入自动分析与抗分析保护互相抬升的新阶段。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">问题也不是“要不要让 AI 参与安全工作”，而是“当 AI 已经开始参与安全工作时，企业、安全人员与行业有没有完成相应的重构”。对企业来说，关键在于流程、权限、审计与补丁治理；对安全人员来说，关键在于理解、验证、编排与责任能力；对行业来说，关键在于建立更成熟的评测、披露与治理基础设施。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">如果把全文进一步收束到一条更底层的判断，那么 AI 时代真正普遍化的，不应只是“人人都能调用更强工具”，而应是 “人人都必须学会更清楚地定义问题、更严格地设置约束、更认真地验证结果”。模型降低了执行门槛，却没有取消问题定义的成本，也没有替人承担结果验证的责任。无论是企业流程重写、漏洞赏金报告质量失控，还是逆向和利用分析中的误判累积，最后暴露出来的都不是同一个模型够不够强，而是使用者是否具备把任务说清楚、把结果查明白的能力。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">还要看到，模型能力始终有上限，知识却没有上限。模型可以更新，可以更强，但它并不会自动替代领域专家长期积累下来的经验知识。真正决定攻防质量的，仍然是对系统、业务、漏洞成因、利用条件、防守约束和历史案例的深度理解。模型越强，这种知识的重要性反而越高，因为模型只能放大已有知识结构，无法凭空替代知识本身。谁拥有更深的领域理解，谁就更能把 AI 变成生产力；谁缺少知识支撑，谁就更容易把 AI 变成看似高效、实则失真的噪音放大器。</span></p><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 2em;"><span leaf="">如果借用历史经验来概括这一变化，那么 AI 之于安全攻防，或许并不只是一次更高效的工具升级，而更像一次迫使规则、基础设施与职业分工重新调整的系统性拐点。下一轮攻防门槛最终会由谁决定，不在于谁最早喊出“AI 安全”，而在于谁最早完成了从工具能力到组织能力的整套迁移。</span></p></div><div style="text-align: left;justify-content: flex-start;display: flex;flex-flow: row;margin: 10px 0px;box-sizing: border-box;"><div style="display: inline-block;width: 100%;vertical-align: top;align-self: flex-start;flex: 0 0 auto;border-style: solid;border-width: 0px 0px 2px;border-bottom-color: rgb(161, 218, 245);box-sizing: border-box;"><div style="justify-content: flex-start;display: flex;flex-flow: row;margin: 0px 0px 3px;box-sizing: border-box;"><div style="display: inline-block;vertical-align: middle;width: auto;align-self: center;flex: 0 0 auto;min-width: 5%;max-width: 100%;height: auto;padding: 0px 0px 0px 8px;box-sizing: border-box;"><div style="text-align: justify;color: rgb(55, 100, 139);box-sizing: border-box;"><p style="white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;"><strong style="box-sizing: border-box;"><span leaf="">参考文献</span></strong></p></div></div></div></div></div><div style="text-align: center;box-sizing: border-box;"><div style="display: inline-block;width: 100%;height: 240px;vertical-align: top;overflow-y: auto;box-sizing: border-box;"><p style="font-size: 12px;text-align: left;box-sizing: border-box;"><ol style="list-style-type: decimal;box-sizing: border-box;padding-left: 20px;list-style-position: outside;" class="list-paddingleft-2"><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，Project Naptime: Evaluating      Offensive Security Capabilities of Large Language Models，2024-06-20。<a href="https://projectzero.google/2024/06/project-naptime.html" target="_blank">https://projectzero.google/2024/06/project-naptime.html</a></span></p></li><li style="box-sizing: border-box;"><p style="text-align: justify;white-space: normal;margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google Project Zero，From Naptime to Big Sleep: Using      Large Language Models To Catch Vulnerabilities In Real-World Code，2024-11-01。<a href="https://projectzero.google/2024/10/from-naptime-to-big-sleep.html" target="_blank">https://projectzero.google/2024/10/from-naptime-to-big-sleep.html</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Google，A summer of security: empowering cyber defenders      with AI，2025-07-15。<a href="https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/" target="_blank">https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Project Glasswing: Securing critical software for      the AI era，2026-04-07。<a href="https://www.anthropic.com/glasswing" target="_blank">https://www.anthropic.com/glasswing</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude Mythos Preview，2026-04-07。<a href="https://red.anthropic.com/2026/mythos-preview/" target="_blank">https://red.anthropic.com/2026/mythos-preview/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Occupational changes during the      20th century，2006-03。<a href="https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm" target="_blank">https://www.bls.gov/opub/mlr/2006/article/occupational-changes-during-the-20th-century.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Federal Highway Administration，The Evolution of MUTCD。<a href="https://mutcd.fhwa.dot.gov/kno-history.htm" target="_blank">https://mutcd.fhwa.dot.gov/kno-history.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">腾讯云黑客松官网；腾讯云开发者社区《400+极客菁英共聚羊城，见证国内首个AI智能渗透挑战赛》。<a href="https://tch.cloud.tencent.com/" target="_blank">https://tch.cloud.tencent.com/</a> ；<a href="https://cloud.tencent.com/developer/article/2651925" target="_blank">https://cloud.tencent.com/developer/article/2651925</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，AIxCC:      AI Cyber Challenge；AI Cyber Challenge marks      pivotal inflection point for cyber defense。<a href="https://www.darpa.mil/research/programs/ai-cyber" target="_blank">https://www.darpa.mil/research/programs/ai-cyber</a> ；<a href="https://www.darpa.mil/news/2025/aixcc-results" target="_blank">https://www.darpa.mil/news/2025/aixcc-results</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Defense Advanced Research Projects      Agency (DARPA)，Cyber      Grand Challenge (CGC)。<a href="https://www.darpa.mil/research/programs/cyber-grand-challenge" target="_blank">https://www.darpa.mil/research/programs/cyber-grand-challenge</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Claude is competitive with humans      in (some) cyber competitions，2025-08-09。<a href="https://red.anthropic.com/2025/cyber-competitions/" target="_blank">https://red.anthropic.com/2025/cyber-competitions/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">National Institute of Standards and      Technology (NIST)，Artificial      Intelligence Risk Management Framework: Generative Artificial Intelligence      Profile，2024-07-26。<a href="https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence" target="_blank">https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Open Worldwide Application Security      Project (OWASP) GenAI Security Project，OWASP Top 10 for LLM is now the GenAI Security      Project and promoted to OWASP Flagship status，2025-03-26。<a href="https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/" target="_blank">https://genai.owasp.org/2025/03/26/project-owasp-promotes-genai-security-project-to-flagship-status/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic Frontier Red Team，Partnering with Mozilla to      improve Firefox&#39;s security，2026-03-06。<a href="https://red.anthropic.com/2026/firefox/" target="_blank">https://red.anthropic.com/2026/firefox/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Hanzhuo Tan, Qi Luo, Jing Li, Yuqun      Zhang，LLM4Decompile:      Decompiling Binary Code with Large Language Models，EMNLP 2024。<a href="https://aclanthology.org/2024.emnlp-main.203/" target="_blank">https://aclanthology.org/2024.emnlp-main.203/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anton Tkachenko, Dmitrij Suskevic,      Benjamin Adolphi，Deconstructing      Obfuscation: A four-dimensional framework for evaluating Large Language      Models assembly code deobfuscation capabilities，2025。<a href="https://arxiv.org/abs/2505.19887" target="_blank">https://arxiv.org/abs/2505.19887</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Abraham Clements et al.，Towards LLM-Resistant Software      Protection: Agent Failure Patterns in CTF Reverse Engineering，NDSS BAR 2026。<a href="https://www.ndss-symposium.org/ndss-paper/auto-draft-657/" target="_blank">https://www.ndss-symposium.org/ndss-paper/auto-draft-657/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Cybench: A Framework for Evaluating      Cybersecurity Capabilities and Risks of Language Models。<a href="https://cybench.github.io/" target="_blank">https://cybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">BountyBench: Dollar Impact of AI Agent      Attackers and Defenders on Real-World Cybersecurity Systems。<a href="https://bountybench.github.io/" target="_blank">https://bountybench.github.io/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Anthropic，Coordinated vulnerability disclosure for      Claude-discovered vulnerabilities，最后更新于 2026-03-06。<a href="https://www.anthropic.com/coordinated-vulnerability-disclosure" target="_blank">https://www.anthropic.com/coordinated-vulnerability-disclosure</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Node.js Project，New HackerOne Signal Requirement      for Vulnerability Reports，最后更新于 2026-02-19。<a href="https://nodejs.org/en/blog/announcements/hackerone-signal-requirement" target="_blank">https://nodejs.org/en/blog/announcements/hackerone-signal-requirement</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">HackerOne，Code of Conduct。<a href="https://www.hackerone.com/policies/code-of-conduct" target="_blank">https://www.hackerone.com/policies/code-of-conduct</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Daniel Stenberg，The end of the curl bug-bounty，2026-01-26。<a href="https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/" target="_blank">https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">前沿观察 赛事纪实：从腾讯云黑客松，洞见智能体时代的攻防新格局，2026-04-17。<a href="https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q" target="_blank">https://mp.weixin.qq.com/s/f94uaYgqiSSx-3Vz0kP4_Q</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">U.S. Bureau of Labor Statistics，Changes in the U.S. occupational      mix from 1860 to 2015，2019-08。<a href="https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm" target="_blank">https://www.bls.gov/opub/mlr/2019/beyond-bls/changes-in-the-us-occupational-mix-from-1860-to-2015.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Martin Fiszbein, Jeanne Lafortune, Ethan      G. Lewis, José Tessada，Powering Up Productivity: The Effects of Electrification on      U.S. Manufacturing，National Bureau of Economic      Research (NBER) Working Paper 28076，2020；2024-04 修订。<a href="https://www.nber.org/papers/w28076" target="_blank">https://www.nber.org/papers/w28076</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">James Bessen，Toil and Technology，International Monetary Fund (IMF) Finance &amp; Development，2015-03。<a href="https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm" target="_blank">https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm</a></span></p></li><li style="box-sizing: border-box;"><p style="margin: 0px;padding: 0px;box-sizing: border-box;text-indent: 0px;"><span leaf="">Timothy F. Bresnahan, Manuel Trajtenberg，General Purpose Technologies      &#34;Engines of Growth?&#34;，National Bureau of      Economic Research (NBER) Working Paper 4148，1992-08。<a href="https://www.nber.org/papers/w4148" target="_blank">https://www.nber.org/papers/w4148</a></span></p></li></ol></p></div></div></div><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box;overflow-wrap: break-word !important;clear: both;min-height: 1em;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;" data-pm-slice="6 2 []"><strong data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 16px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;letter-spacing: 0.544px;display: inline;color: rgb(62, 62, 62);font-family: 楷体;text-align: left;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;letter-spacing: normal;color: rgb(255, 0, 0);font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><strong style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;display: inline;color: rgb(255, 79, 121);font-size: 16px;"><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;font-size: 17px;text-decoration-style: solid;text-decoration-color: rgb(255, 0, 0);"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">本公众号发布、转载的文章所涉及的技术、思路、工具仅供学习交流，任何人不得将其用于非法用途及盈利等目的，否则后果自行承担！</span></span></strong></span></strong></p><p style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;text-align: center;"><span data-pm-slice="0 0 []" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">点这里 <img data-aistatus="1" alt="图片" class="rich_pages wxw-img __bg_gif" data-ratio="0.5982532751091703" data-w="458" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;vertical-align: middle;height: auto !important;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;font-family: PingFangSC-Regular, &#34;PingFang SC&#34;;text-indent: 28px;color: rgb(62, 62, 62);font-size: 16px;width: 61.9922px !important;visibility: visible !important;" data-width="100%" data-imgfileid="100041710" src="https://wechat2rss.xlab.app/img-proxy/?k=6849f525&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_gif%2FMfTd6rd9CyvNRMW8I9cvI1CK5gKiaYqg2veTn9t9dAe1GxYic7pAvgvRIKNFickConFyX8AvW2reAq8GchJI6aBpA%2F640%3Fwx_fmt%3Dgif%26wxfrom%3D5%26wx_lazy%3D1%26tp%3Dwebp%23imgIndex%3D14"/></span><span style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgb(127, 229, 230);font-family: &#34;Helvetica Neue&#34;, Helvetica, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 14px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.578px;orphans: 2;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;float: none;display: inline !important;"><span leaf="" data-pm-slice="1 1 [&#34;para&#34;,null,&#34;node&#34;,{&#34;tagName&#34;:&#34;span&#34;,&#34;attributes&#34;:{&#34;style&#34;:&#34;color: rgb(127, 229, 230); font-family: \&#34;Helvetica Neue\&#34;, Helvetica, \&#34;Hiragino Sans GB\&#34;, \&#34;Microsoft YaHei\&#34;, Arial, sans-serif; font-size: 14px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: 0.578px; orphans: 2; text-align: center; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px;  background-color: rgb(255, 255, 255); text-decoration-thickness: initial; text-decoration-style: initial; text-decoration-color: initial; display: inline !important; float: none;&#34;},&#34;namespaceURI&#34;:&#34;http://www.w3.org/1999/xhtml&#34;}]" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;">关注我们，一键三连～</span></span></span></p><p class="mp_profile_iframe_wrp" nodeleaf="" style="-webkit-tap-highlight-color: rgba(0, 0, 0, 0);margin: 0px 0px 24px;padding: 0px;outline: 0px;max-width: 100%;box-sizing: border-box !important;overflow-wrap: break-word !important;color: rgba(0, 0, 0, 0.9);font-family: &#34;PingFang SC NEW&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;font-size: 17px;font-style: normal;font-variant-ligatures: normal;font-variant-caps: normal;font-weight: 400;letter-spacing: 0.544px;orphans: 2;text-align: justify;text-indent: 0px;text-transform: none;widows: 2;word-spacing: 0px;-webkit-text-stroke-width: 0px;white-space: normal;background-color: rgb(255, 255, 255);text-decoration-thickness: initial;text-decoration-style: initial;text-decoration-color: initial;"><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe js_wx_tap_highlight" data-pluginname="mpprofile" data-nickname="华为安全应急响应中心" data-alias="HUAWEI_PSIRT" data-index="0" data-from="2" data-headimg="http://mmbiz.qpic.cn/sz_mmbiz_png/Pf9eicDVDMxHbPW1POGs9HHQCUGXBXg7u6TCtI2ab5DdIEfxJWcR46krXgudVuibfibsqRYlAtN2RLdaiaOCosQMSw/300?wx_fmt=png&amp;wxfrom=19" data-signature="华为安全应急响应中心（HUAWEI PSIRT）官方公众号。" data-id="MzI0MTY5NDQyMw==" data-is_biz_ban="0" data-origin_num="49" data-biz_account_status="0" data-service_type="1" data-verify_status="2"></mp-common-profile></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=9e022ddc&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486092%26idx%3D3%26sn%3D1fdedc10e9db976f0d01304dd4490da7">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Thu, 09 Jul 2026 18:02:00 +0800</pubDate>
    </item>
    <item>
      <title>AI 时代下软件保护对抗的现状与趋势</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486086&amp;idx=1&amp;sn=dc36e3f24efce79c417db97b1b01cd31</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>漏洞战争</span> <span>2026-06-28 07:55</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=37bd5bf3&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2syIZx0xljua4dht69zNDG6DVKaNszyVU6jT91PaVeUBs0Z7XmibRoWrYfKtmNs8Thoz5iakKjA0Q66icadmBKjibTYSnQibI6e1S8ug%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <h1 style="color: #2B77BF;text-align: center;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="0"><span leaf="">一、核心结论</span></h1><p data-layout-id="2" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 正在改变软件保护的攻防结构，但还没有让代码混淆、App Shielding、运行时保护失效。Promon 的 Q1 2026 报告给出的关键判断是：领先 LLM 已经具备一定反混淆能力，可以读取反汇编代码、分析程序逻辑并尝试恢复原始代码；但在真实移动 App 场景中，尤其是 ARM 架构和多层混淆叠加后，AI 的成功率仍会显著下降。</span></span></p><p data-layout-id="3" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这意味着软件保护进入了一个新阶段：过去主要对抗人类逆向工程师的理解成本，现在还必须对抗 AI Agent 的自动化分析、批量试错和工具编排能力。保护目标不再只是“让人看不懂”，而是让 AI 的自动化破解链路变慢、变错、变贵，并迫使攻击者回到人工验证。</span></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="4"><span leaf="">二、Promon 报告揭示的现状</span></h1><h2 style="color: #2B77BF;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="5"><span leaf="">1. AI 已经能攻击混淆代码，但能力高度受场景限制</span></h2><p data-layout-id="6" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 测试了 10 个领先 AI 模型，对 OLLVM 混淆代码进行反混淆评估，覆盖 x86 与 ARM 两种架构，以及 raw assembly 和 Ghidra pseudocode 两种输入形态。结果显示，AI 对混淆代码确实构成现实威胁，但不是无条件成功。</span></span></p><p data-layout-id="7" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">在最强的三层混淆 SUB（指令替代） + FLA（控制流平坦化） + BCF（虚假控制流） 下，最强模型在 x86 架构上仍能以约 20%–36% 的成功率恢复可工作代码；但在 ARM 场景下，平均成功率降到 8.5%，多数模型只有个位数成功率。Promon 因此认为，ARM 移动 App 当前对 AI 反混淆有更强抵抗力。</span></span></p><p style="text-align: center;" nodeleaf=""><img class="rich_pages wxw-img" data-aistatus="1" data-imgfileid="100002432" data-ratio="0.4537037037037037" data-s="300,640" type="block" data-type="png" data-w="1080" src="https://wechat2rss.xlab.app/img-proxy/?k=325b3748&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2syBOAGiaDLHDGecXBzkYQYwesnZcX9MPWC6cLNgb90CXrCXFqV2OlYXl2zFfeKEwChYIH8Czxv6wdIDUd6rib53b71jTKRotIL2E%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></p><p data-layout-id="8" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这对移动 App 保护是一个重要信号：AI 对软件保护的威胁已经真实存在，但现阶段仍强烈依赖架构、输入质量、混淆层数和模型能力。</span></span></p><h2 data-layout-id="9" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">2. 分层混淆仍然有效，而且效果不是简单相加</span></h2><p data-layout-id="10" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">报告最重要的工程结论之一是：多种混淆技术组合后会产生乘法效应，而不是线性叠加。</span></span></p><p data-layout-id="11" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 特别指出，Control Flow Flattening 与 Bogus Control Flow 组合时，会显著放大代码结构复杂度。FLA 会生成中心调度器，BCF 再在这些调度点插入虚假分支，使控制流复杂度被放大。报告给出的数据是：相较单独 BCF，组合后复杂度在 x86 上放大 4.18x，在 ARM 上放大 5.50x。</span></span></p><p style="text-align: center;" nodeleaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.5666666666666667" data-s="300,640" data-type="png" data-w="1080" type="block" data-imgfileid="100002433" src="https://wechat2rss.xlab.app/img-proxy/?k=692fc2dc&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2szMXZcbicB1F0Ivwpx22NicPmBcTXDMwsGiczOuDuCibLqdAvynibuibXrYmtT7OhekW7XM8xK0yZjwxo22zKYib4xNgG1y9VDayrjkiaU%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></p><p data-layout-id="12" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这说明混淆并没有因为 AI 出现而失效。恰恰相反，面对 AI 自动化逆向，单点混淆会显得薄弱，而组合式、分层式混淆成为基本要求。</span></span></p><h2 data-layout-id="13" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">3. AI 工具有内生错误率，防守方可以利用这个弱点</span></h2><p data-layout-id="14" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 一个很有启发性的发现是：AI 即使处理未混淆代码，也不是 100% 准确。报告称，没有模型在 clean code 上超过 86% 成功率；GPT-4o 在 clean x86 assembly 上只有 48% 成功率；clean ARM assembly 的平均成功率也只有 63.7%。</span></span></p><p data-layout-id="15" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这改变了防守方对“AI 破解能力”的理解。攻击者并不是从 100% 准确率开始，再被混淆削弱；攻击者一开始就会产生大量错误输出。混淆的作用是把这种错误率进一步放大。</span></span></p><p data-layout-id="16" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">在真实攻击中，AI Agent 会批量处理一个 App 中成百上千个函数。只要输出中混入大量错误函数，攻击者就很难自动判断哪些恢复结果可信、哪些会导致逻辑错误。于是，攻击流程会被迫进入人工验证，自动化优势被削弱。</span></span></p><h2 data-layout-id="17" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">4. 反编译器质量会直接决定 AI 攻击效果</span></h2><p data-layout-id="18" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 报告指出，AI 拿到 raw assembly 和拿到 Ghidra 生成的 pseudocode，成功率差异很大。以三层 ARM 混淆为例，Claude Opus 4.5 从 pseudocode 输入可达到 50% 成功率，但从 raw ARM assembly 只有 24%；GPT-4o 则是 10% 对 2%。</span></span></p><p data-layout-id="19" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这说明 AI 不是独立完成逆向。它高度依赖上游工具提供的输入质量。反编译器越能生成清晰、结构化、接近 C 语言的伪代码，AI 越容易理解逻辑；反编译器输出越破碎、误导、类型错误、函数边界错误，AI 越容易产生错误推理。</span></span></p><p data-layout-id="20" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这也是未来软件保护的重要方向：防守不只要混淆二进制本身，还要主动降低反编译器输出质量，也就是 anti-decompilation。</span></span></p><h2 data-layout-id="21" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">5. 不同 AI 模型的威胁等级差异很大</span></h2><p data-layout-id="22" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 测试发现，模型能力差距明显。顶级模型在高强度混淆下仍保持一定成功率，而部分模型在三层 ARM assembly 场景下几乎失败。</span></span></p><p data-layout-id="23" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">因此，软件保护不能只看平均模型能力。对普通自动化攻击，可以参考平均成功率；但对金融、支付、游戏、流媒体、SDK、版权保护、企业软件和高价值算法，应该按顶级模型能力建模。攻击者一旦有足够收益，通常会选择更强模型、更好的反编译器和更完整的自动化流水线。</span></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="24"><span leaf="">三、AI 时代软件保护的本质变化</span></h1><h2 style="color: #2B77BF;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="25"><span leaf="">1. 对手从“人类逆向者”变成“AI Agent 破解流水线”</span></h2><p data-layout-id="26" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">传统软件保护假设攻击者是人类：打开 IDA、Ghidra、x64dbg、Frida，逐步理解关键函数，定位授权校验、反调试、Hook 检测、协议签名和业务限制点。防守方的策略是提高人的阅读成本、定位成本和修改成本。</span></span></p><p data-layout-id="27" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 时代的对手模型不同。更现实的攻击者会构建一条流水线：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="28" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">用反编译器批量生成 pseudocode。</span></span></p></li><li><p data-layout-id="29" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">用 LLM 总结函数语义、识别可疑路径。</span></span></p></li><li><p data-layout-id="30" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">用脚本和调试器验证 Hook 点。</span></span></p></li><li><p data-layout-id="31" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">用 Agent 记录失败原因并规划下一轮尝试。</span></span></p></li><li><p data-layout-id="32" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">用自动化测试环境验证补丁是否生效。</span></span></p></li><li><p data-layout-id="33" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">最后由人类确认关键路径和可变现方式。</span></span></p></li></ul><p data-layout-id="34" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">人类不再逐行看代码，而是管理 AI Agent 的任务、结果和验证。这会显著降低攻击成本。</span></span></p><h2 data-layout-id="35" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">2. 保护目标从“隐藏逻辑”转向“破坏自动化闭环”</span></h2><p data-layout-id="36" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">过去软件保护常追求“隐藏关键逻辑”。但 AI Agent 攻击更依赖闭环：观察输入、生成假设、修改或 Hook、运行验证、根据反馈调整策略。</span></span></p><p data-layout-id="37" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">因此，未来软件保护的目标应包括：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="38" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">降低 AI 可读输入质量。</span></span></p></li><li><p data-layout-id="39" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">增加错误恢复结果的比例。</span></span></p></li><li><p data-layout-id="40" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">减少高信号反馈。</span></span></p></li><li><p data-layout-id="41" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">延迟关键校验结果。</span></span></p></li><li><p data-layout-id="42" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">让单点补丁无法证明成功。</span></span></p></li><li><p data-layout-id="43" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">把关键判断移到服务端和业务风控侧。</span></span></p></li><li><p data-layout-id="44" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">让攻击者必须跨模块、跨会话、跨设备、跨账号验证。</span></span></p></li></ul><p data-layout-id="45" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">换句话说，保护不是追求绝对不可逆，而是让 AI Agent 无法低成本自动完成“理解-修改-验证-规模化”的闭环。</span></span></p><h2 data-layout-id="46" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">3. 客户端保护必须和服务端可信结合</span></h2><p data-layout-id="47" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 会让客户端逆向更高效，所以高价值逻辑不能只依赖客户端隐藏。授权、权益、支付、风控、内容访问、游戏经济、版权下载、SDK 调用资格等关键判断，应尽可能与服务端状态绑定。</span></span></p><p data-layout-id="48" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">更稳健的模式是：</span></span></p><p data-layout-id="49" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">客户端做混淆、完整性校验、反调试、反 Hook、反重打包、Root/Jailbreak 检测。</span></span></p><p data-layout-id="50" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">服务端做 App attestation、设备可信、账号风险、行为序列、授权状态和异常收益检测。</span></span></p><p data-layout-id="51" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">客户端保护负责提高攻击成本，服务端保护负责判断请求是否可信。</span></span></p><p data-layout-id="52" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 时代的软件保护不应只看“二进制有没有被破解”，还要看“破解后的客户端能否持续获得业务收益”。</span></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="53"><span leaf="">四、未来趋势判断</span></h1><h2 style="color: #2B77BF;font-size: 17px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;" data-layout-id="54"><span leaf="">趋势 1：AI 反混淆能力会持续提升，ARM 当前优势不会永久存在</span></h2><p data-layout-id="55" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 认为 ARM 当前比 x86 更抗 AI，一个重要原因是训练数据差异。公开互联网中 x86 反汇编、逆向文章、恶意软件分析资料更多，而 ARM 高质量逆向语料相对少。</span></span></p><p data-layout-id="56" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">但这个差距可能缩小。随着移动逆向、IoT、车载、边缘设备、ARM 服务器资料增多，模型对 ARM 的理解能力会提升。防守方不能把当前 ARM 优势当作长期护城河。</span></span></p><h2 data-layout-id="57" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">趋势 2：Anti-decompilation 会成为软件保护关键方向</span></h2><p data-layout-id="58" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 报告显示，pseudocode 输入能显著提升 AI 成功率。因此，未来保护不会只做传统控制流混淆，还会更多针对反编译器：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="59" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">破坏函数边界识别。</span></span></p></li><li><p data-layout-id="60" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">制造误导性类型信息。</span></span></p></li><li><p data-layout-id="61" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">让伪代码结构不稳定。</span></span></p></li><li><p data-layout-id="62" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">增加不可达但看似关键的路径。</span></span></p></li><li><p data-layout-id="63" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">让反编译输出语义与真实执行语义偏离。</span></span></p></li></ul><p data-layout-id="64" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这类保护的价值在于，它攻击的是 AI Agent 的上游输入质量。</span></span></p><h2 data-layout-id="65" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">趋势 3：软件保护会从静态保护转向动态、分层和服务端协同</span></h2><p data-layout-id="66" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">单一混淆很难长期对抗 AI。未来更有效的是组合防护：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="67" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">静态层：控制流混淆、指令替换、字符串加密、虚拟化保护。</span></span></p></li><li><p data-layout-id="68" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">反分析层：反调试、反 Hook、反模拟器、反 Frida、反重打包。</span></span></p></li><li><p data-layout-id="69" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">完整性层：运行时校验、代码段校验、资源校验、签名校验。</span></span></p></li><li><p data-layout-id="70" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">可信层：App attestation、设备绑定、密钥保护、远程证明。</span></span></p></li><li><p data-layout-id="71" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">业务层：账号风控、行为异常、权益校验、收益异常检测。</span></span></p></li></ul><p data-layout-id="72" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 迫使软件保护从“代码工程”升级为“代码 + 运行时 + 设备 + 服务端 + 业务”的体系工程。</span></span></p><h2 data-layout-id="73" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">趋势 4：保护评估会从人工渗透测试转向 AI 对抗基准测试</span></h2><p data-layout-id="74" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon 的方法本身也代表一个趋势：未来安全团队会定期用不同模型、不同反编译器、不同混淆配置测试保护效果。</span></span></p><p data-layout-id="75" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">评估指标也会变化：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="76" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 能否恢复可执行逻辑。</span></span></p></li><li><p data-layout-id="77" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">错误恢复比例是多少。</span></span></p></li><li><p data-layout-id="78" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 能否定位关键授权路径。</span></span></p></li><li><p data-layout-id="79" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 是否能生成有效 Hook 假设。</span></span></p></li><li><p data-layout-id="80" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI Agent 是否能完成自动验证闭环。</span></span></p></li><li><p data-layout-id="81" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">保护是否能让攻击者必须人工介入。</span></span></p></li></ul><p data-layout-id="82" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">软件保护不再只问“能不能被人逆向”，而是问“顶级 AI Agent 在多少成本内能否规模化破解”。</span></span></p><h2 data-layout-id="83" style="font-size: 17px;font-weight: 500;color: #2B77BF;line-height: 1.8;margin-bottom: 12px;"><span leaf="">趋势 5：防护设计会主动利用 AI 的弱点</span></h2><p data-layout-id="84" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">未来有效的软件保护会更有意识地针对 AI Agent 的弱点设计：</span></span></p><ul style="list-style-type: square;" class="list-paddingleft-1"><li><p data-layout-id="85" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对上下文窗口有限：增加跨函数、跨模块、跨会话依赖。</span></span></p></li><li><p data-layout-id="86" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对模型依赖高质量输入：降低反编译器输出质量。</span></span></p></li><li><p data-layout-id="87" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对模型容易错误归因：设计诱饵路径和低信号失败。</span></span></p></li><li><p data-layout-id="88" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对 Agent 需要验证反馈：延迟、模糊或分散关键反馈。</span></span></p></li><li><p data-layout-id="89" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对批量自动化：让错误输出难以被自动筛除。</span></span></p></li><li><p data-layout-id="90" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">针对工具链依赖：检测调试器、Hook 框架、模拟器和重打包环境。</span></span></p></li></ul><p data-layout-id="91" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">这不是“骗 AI”本身，而是系统性提高 AI 自动化攻击的成本。</span></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="92"><span leaf="">五、对企业和产品团队的启示</span></h1><p style="text-align: center;" nodeleaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="1.3953703703703704" data-s="300,640" data-type="png" data-w="1080" type="block" data-imgfileid="100002434" src="https://wechat2rss.xlab.app/img-proxy/?k=0bd85c2d&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2sxiacaTsEzibgc0jfMoN8drrFlxp1b5xb9Ul9pps8MQMU99wuPT2jCsztTmeuiaUKpUic1mS8ia7b4bSbStOibeGdTtakWSCCk9qWIBw%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></p><p data-layout-id="93" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第一，把 AI 反混淆纳入威胁模型。只按传统人工逆向能力设计保护，已经低估风险。</span></span></p><p data-layout-id="94" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第二，把三层或多层混淆作为敏感逻辑保护基线。单点混淆在 AI Agent 面前容易被批量试探。</span></span></p><p data-layout-id="95" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第三，移动 App 当前相对更有优势，但不要依赖 ARM 架构红利。模型训练数据会补齐，工具链也会进步。</span></span></p><p data-layout-id="96" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第四，投入 anti-decompilation。AI 越依赖反编译器，破坏 pseudocode 质量越有战略价值。</span></span></p><p data-layout-id="97" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第五，把客户端保护和服务端可信绑定。客户端负责抬高破解成本，服务端负责判断请求可信度和业务结果合理性。</span></span></p><p data-layout-id="98" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">第六，建立持续评估机制。每次模型能力、反编译器版本、保护配置变化，都可能改变攻防平衡。</span></span></p><h1 style="color: #2B77BF;font-size: 20px;font-weight: 500;line-height: 1.8;margin-bottom: 12px;text-align: center;" data-layout-id="99"><span leaf="">六、结语</span></h1><p data-layout-id="100" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">Promon Q1 2026 报告的真正价值，不是证明“AI 已经破解混淆”，也不是证明“混淆仍然安全”，而是把软件保护带入了可量化的 AI 对抗阶段。</span></span></p><p data-layout-id="101" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">AI 的出现改变了攻击成本结构。攻击者可以更快理解代码、更快生成假设、更快批量验证。防守方也必须改变目标：不再只对抗人类阅读，而是对抗 AI Agent 的自动化逆向流水线。</span></span></p><p data-layout-id="102" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="color: rgba(0, 0, 0, 0.9);">未来软件保护的竞争点，将不只是混淆强度，而是能否持续破坏 AI 的输入质量、推理稳定性、验证闭环和规模化能力。能做到这一点的软件保护，才是真正面向 AI 时代的保护。</span></span></p><h2 style="font-size: 17px;font-weight: 500;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 12px;" data-pm-slice="2 4 []"><span leaf="">参考来源</span></h2><ul style="font-size: 15px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" class="list-paddingleft-1"><li style="margin-bottom: 0px;"><p><span leaf="">Promon, 2026, App Threat Report 2026 Q1: The State of Code Obfuscation Against AI</span></p></li><li style="margin-bottom: 0px;"><p><span leaf="">Promon, 2025, AI powered mobile app attacks: What app shielding can and can&#39;t stop</span></p></li></ul><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=86cfbf48&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486086%26idx%3D1%26sn%3Ddc36e3f24efce79c417db97b1b01cd31">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sun, 28 Jun 2026 07:55:00 +0800</pubDate>
    </item>
    <item>
      <title>用 GPT-5.4 单挑 NCTF 团队赛，成功解出91.7%的题目</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486075&amp;idx=1&amp;sn=5c8c4a448349149a72daf771e643ce93</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>漏洞战争</span> <span>2026-04-06 10:59</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=0ed82de7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2sxPfTC95ZZM7LjuxVyAz0I8WOWbj2sDBpe7oqPicG6Zhdw9vgFh19cPT3KN4uJZ7aib0BtJlPzozs6DicSwh5WibDZicMxSptzlbtj0%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p data-layout-id="0" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">自从买了token套餐之后，每天不把token用完就有点焦虑。于是，放假这2天，就打算用GPT-5.4来打CTF比赛。网上找了下，刚好南京邮电大学在举办NCTF 2026比赛，就拿来作实验，看一个人带着GPT-5.4，如何单挑整个团队赛（4人赛）。</span></p><p data-layout-id="1" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">刚才9点（4月6日）的时候，比赛已结束。</span></p><p data-layout-id="2" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">最终成绩：<span textstyle="" style="font-weight: bold;">24道题，成功解出22道，解题率91.7%</span>。</span></p><p class="mp_profile_iframe_wrp" nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="0" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><p data-layout-id="2" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">排名34，共有参赛队伍915支，有得分的433支队伍。</span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="3"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img js_insertlocalimg" data-ratio="0.26380368098159507" data-s="300,640" data-type="png" data-w="652" type="block" data-imgfileid="100002419" src="https://wechat2rss.xlab.app/img-proxy/?k=716f6963&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2sw4mk6aAf24sXZIRibKpQx4I6ksudZIRia8z4Nv8fmM27bJLmE1IDialMt2dZhhyhlNz5O27U3N6BIPgeTOpNS7Le4ms0rmoRotjA%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><p data-layout-id="4" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在这场所谓的“技术对决”中，我没有写一行代码，没有做任何手动分析，甚至连IDA、JADX这些最基本的反编译工具都没装。我不装任何MCP，不给任何技术指导，我在这场比赛中的唯一身份是——“题目的搬运工”，最多在任务失败时，让它再重试下。</span></p><p data-layout-id="5" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">别问我为什么没用 claude， 因为穷。今天，可以聊聊这场实验背后的细节，以及它对当前安全行业释放的信号。</span></p><h1 data-layout-id="6" style="font-size: 20px;font-weight: 500;color: rgba(43, 119, 191, 1);line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf=""><span textstyle="" style="font-weight: bold;">01 极致的“躺平”：我是如何打这场比赛的？</span></span></h1><p data-layout-id="7" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我的武器库极其简单：<span textstyle="" style="font-weight: bold;">Codex + GPT-5.4</span>以及<span textstyle="" style="font-weight: bold;">Trae + GPT-5.4</span>。</span></p><p data-layout-id="8" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">我的工作流可以用“三步走”概括：</span></p><p data-layout-id="9" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">搬运</span>：把题目描述、附件原封不动地扔给AI。容器有启动时长限制，有时超时会重启换端口，这个需要再告诉下AI。</span></p><p data-layout-id="10" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">装死</span>：绝对不给任何“你可以试试看XX算法”、“这里有个XX漏洞”的提示，完全不引导。</span></p><p data-layout-id="11" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">重试</span>：当AI报错或解不出时，我的回复只有三类：“重试”、“换个思路再试下”、“这么简单你都做不出来？再想想”。</span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="12"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.15376106194690264" data-s="300,640" data-type="png" data-w="904" type="block" data-imgfileid="100002422" src="https://wechat2rss.xlab.app/img-proxy/?k=0f63500a&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2syzBOyzbzOicwhL4r2SXST1EpRCYgufbNPDB4aEFpcepIVGTxOiar4v40vhCV2VFxIZgR0kxmzgI85E2XiaOUEwr8zzJPSAx20Sy4%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><p data-layout-id="13" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">除agent自带工具外，不再提供任何工具，也没有手工搭建环境（全靠AI在沙盒里自己搞），遇到二进制文件和APK，全靠AI自己找工具逆向，反汇编它会用objdump，apk逆向会安装baksmali与Androguard，也会自动gdb调试。在失败中不断让AI自我反思、自我迭代，直到把Flag吐出来。</span></p><p data-layout-id="14" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本地没有的工具就连网搜索，比如盲打后台XSS，自己从网上找webhook.site来接收flag。</span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="15"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.18425925925925926" data-s="300,640" data-type="png" data-w="1080" type="block" data-imgfileid="100002421" src="https://wechat2rss.xlab.app/img-proxy/?k=ea11078a&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2sw07ZnNh0affvct5z1dIAKwiaOZL9vX1r3chHONjR3B6Wvexru5VaUYl8DG8uPqz64RZXTMD3DNUlKv6LXBnll33n3VOLdiap0wk%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="16"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img js_insertlocalimg" data-ratio="0.48144712430426717" data-s="300,640" data-type="png" data-w="1078" type="block" data-imgfileid="100002420" src="https://wechat2rss.xlab.app/img-proxy/?k=2c741013&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_png%2FtJDT9c8t2swD7REbibTbvwPBdibAQKJFMb5F6iadtXRLq5hjsCTubUwwnCCCr6gicMVNtvelhjFs4NWicUKedbVTuAqTwwH7PLZfGn04hf5L4JB4%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></span></p></div><p data-layout-id="17" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">就这样，比赛还没结束，22道题的Flag就已经躺在我的屏幕上了。</span></p><p data-layout-id="17" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">比赛中1个账号最多只能开2个远程容器实例，如果放开的话，用AI去打将会更快，当然你也可以多建几个账号去开启，也能解决。</span></p><p data-layout-id="18" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">比赛2天，其中一天带娃去商场玩，昨晚又打了一晚麻将，就让AI在家干活：</span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="19"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.562962962962963" data-s="300,640" data-type="jpeg" data-w="1080" type="block" data-imgfileid="100002423" src="https://wechat2rss.xlab.app/img-proxy/?k=be9c6b57&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2swy3o14H1UdqCSGQlbJX0DdTCfAXYxHMfAYM5f7uQfSd0iaOx4qNK5vV8Xd6WicgveqSGduBFZiae19ZArlYicNRicMv7llqhAkvuwU%2F640%3Fwx_fmt%3Djpeg%26from%3Dappmsg"/></span></p></div><p data-layout-id="20" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">手机通过 ToDesk远程控制电脑，看下处理进度，以及延长容器启动时间或提供新IP+端口的变更信息去重试。</span></p><div style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;" data-layout-id="21"><p style="text-align: center;font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.4527777777777778" data-s="300,640" data-type="jpeg" data-w="1080" type="block" data-imgfileid="100002424" src="https://wechat2rss.xlab.app/img-proxy/?k=a8f423d5&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2szrFgnNuIbN22405VLchSiabnVibOTmZEBcej8xZAUCFziadBNjtiak77Ph3GD1He3Fmy120WibpcXAYNC7iaT3kYJRq0ib6MNCgkY1q0%2F640%3Fwx_fmt%3Djpeg%26from%3Dappmsg"/></span></p></div><h1 data-layout-id="22" style="font-size: 20px;font-weight: 500;color: rgba(43, 119, 191, 1);line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf=""><span textstyle="" style="font-weight: bold;">02 工具大PK：同样的GPT-5.4，差距肉眼可见</span></span></h1><p data-layout-id="23" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">在测试过程中，我对比了几个不同的环境，得出的结论非常残酷：</span></p><p data-layout-id="24" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">第一：国产大模型，真的打不过</span></span></p><p data-layout-id="25" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">期间我也尝试用几款主流的国产模型（GLM、Qwen、Kimi）去跑同样的题目，结果搞不出来。很多稍微复杂一点的逻辑绕过、非标准加密、或者长代码的逆向分析，国产模型找不到真正的漏洞点或者算法逆向出现幻觉。在深度的安全攻防推理上，GPT-5.4展现出的逻辑链条完整度，目前国产模型确实难以企及。</span></p><p data-layout-id="26" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf=""><span textstyle="" style="font-weight: bold;">第二：Trae + GPT-5.4 搞不定的，Codex + GPT-5.4 能搞定</span></span></p><p data-layout-id="27" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">手上刚好同时买了gpt和trae，就想设置完全一样的底层模型GPT-5.4进行比较，但两者的解题率却有差异。为什么？</span></p><p data-layout-id="28" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">答案在于Agent工程能力。</span></p><p data-layout-id="29" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">个人感觉Trae在使用体验上要比codex好，但在CTF这种需要“试错-报错-修改环境-再试错”的长链路Agent任务中，它的工具调用、循环反馈、纠错能力要弱于codex，除agent工程能力差异外，可能gpt本身也针对codex作一些适配性训练，使得codex + gpt搭配能达到更好的效果。</span></p><p data-layout-id="30" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">而Codex的Agent调度极其强悍，它能自主搭建本地漏洞环境、自主写脚本编译、自主网上找源码进行现场漏洞挖掘、自主调试Segmentation Fault修改exp，甚至在遇到死胡同时能自己推翻重写。这证明了在AI时代，上层的Agent工程框架，其重要性完全不亚于底层的基座模型。</span></p><h1 data-layout-id="31" style="font-size: 20px;font-weight: 500;color: rgba(43, 119, 191, 1);line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf=""><span textstyle="" style="font-weight: bold;">03 给出题方的“降维打击”：AI时代的出题困境</span></span></h1><p data-layout-id="32" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">站在参赛者的角度，91.7%是个爽文成绩；但站在行业观察者的角度，这反映出当前CTF赛事的一个巨大危机：出题方对AI能力的评估严重不足。</span></p><p data-layout-id="33" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">本次NCTF整体题目难度偏低，完全没有针对AI的“抗性设计”。</span></p><p data-layout-id="34" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">2天的比赛，第1天基本就被人（或者说被AI）做完了。由于AI拉平了个体之间的技术鸿沟，导致各个团队之间根本拉不开差距——以前是你懂PWN我不懂，现在是只要会复制粘贴，大家都是PWN手。</span></p><p data-layout-id="35" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">传统的“套壳题”、“标准算法变种题”、“常规框架漏洞题”，在GPT-5.4面前犹如裸奔。出题人如果还停留在“我把这个点挖深一点、代码混淆厚一点”的传统思路上，注定会被AI轻易秒杀。</span></p><h1 data-layout-id="36" style="font-size: 20px;font-weight: 500;color: rgba(43, 119, 191, 1);line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf=""><span textstyle="" style="font-weight: bold;">04 凛冬已至：安全研究员的生存挑战</span></span></h1><p data-layout-id="37" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">这场实验证明：当一个只会“搬运题目”的人，能靠AI打出91.7%的解题率时，大量初级安全研究员、渗透测试员、甚至部分中级研究员的饭碗，已经在摇摇欲坠了。</span></p><p data-layout-id="38" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">AI对安全行业的影响不是未来式，而是现在进行时。面对这种冲击，我们更应该全面拥抱AI，学会使用它，用AI来解决个人过往搞不定的事情，让自己变强。</span></p><p data-layout-id="39" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">未来的安全研究者，将是那些能够与AI建立&#34;共生关系&#34;的人：<span textstyle="" style="font-weight: bold;">既懂得借助AI突破算力边界，又能在关键节点注入人类独有的直觉、伦理判断和创造性思维。</span></span></p><h1 data-layout-id="40" style="font-size: 20px;font-weight: 500;color: rgba(43, 119, 191, 1);line-height: 1.8;margin-bottom: 12px;text-align: center;"><span leaf=""><span textstyle="" style="font-weight: bold;">写在最后</span></span></h1><p data-layout-id="41" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">24道题解出22道，我并没有感到任何“技术上的成就感”，反而有一种强烈的危机感。</span></p><p data-layout-id="42" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">当安全技术的门槛被大模型彻底踏平，当我们引以为傲的“手搓ROP链”、“逆向硬刚”变成了历史遗迹，我们不禁要问：剥离了工具和代码技巧后，安全研究员最核心的能力到底是什么？</span></p><p data-layout-id="43" style="font-size: 17px;font-weight: 400;color: rgba(0,0,0,0.9);line-height: 1.8;margin-bottom: 24px;"><span leaf="">但玩笑归玩笑，潮水已经涌来，别做那个还在沙滩上用沙子堆城堡的人。</span></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=4cf9fc69&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486075%26idx%3D1%26sn%3D5c8c4a448349149a72daf771e643ce93">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Mon, 06 Apr 2026 10:59:00 +0800</pubDate>
    </item>
    <item>
      <title>别让读书，变成一场“正确”的表演</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486065&amp;idx=1&amp;sn=916323140152e2c7b2ada48050b6a75a</link>
      <description></description>
      <content:encoded><![CDATA[<p>原创 <span>riusksk</span> <span>2026-03-16 22:18</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=0cf283a6&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sxXeCGrUMw5G7A51y74pjyhjPJDnPAFApXulUYEUSH7E0ueclvcRyXTphTib2RI0cqQR9JVECblnU5RiavGsRFxjYn8icYG5LRhDQ%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p>刚和朋友们聊起读书这件事，突然发现一个很有意思的现象：许多人越来越害怕“独自阅读”，反而热衷于把读书变成一场集体仪式。<br/> <br/>有人说，读书本就是很私人的事。只有两个人聊得来、三观合，才有分享的必要；不然，再多的倾诉，根本不是交流，更像在“传教”——硬要大家一起感动、一起共鸣，甚至让彼此的读后感都变得“整齐划一”。<br/> <br/>可现实是，我们总被推着往前走：<br/> <br/>- 明明只想安安静静读完一本书，却要在社交平台打卡、晒进度；<br/>- 明明对某段文字有自己的理解，却要附和“主流解读”，生怕显得“没看懂”；<br/>- 明明更享受独处的阅读时光，却要挤进读书群、参加读书会，仿佛“不社交就不算读书”。<br/> <br/>我们渐渐忘了，读书的本质，从来不是和别人比进度、比见解，而是自己跟自己对话，跟作者对话。<br/> <span style="font-weight: bold;"><br/>一、读书，从来都是“私人的事”</span> <br/>同一本书，有人看到悲悯，有人看到力量，有人看到清醒，有人看到迷茫。这些独一无二的感受，没有对错，也不该有标准答案。<br/> <br/>就像有人读《百年孤独》，看见的是家族轮回的宿命；有人读它，看见的是对抗孤独的勇气。没有谁的理解更“高级”，也没有谁的感受更“正确”——因为每个人的成长经历、心境处境、思考角度，都截然不同。<br/> <br/>读书的意义，从来不是强行统一观点、灌输“标准答案”，而是让你看见相似的热爱与共鸣，也看见各异的理解与碰撞。它让你在文字里找到自己，也让你明白：阅读的自由，就在于允许不同的声音存在。<br/> <br/>真正的阅读，不是“传教”，不是“打卡”，不是“完成KPI”，而是让你在文字里安顿自己，学会独立思考，包容不同的解读，让同频的人因书相遇，因理解而靠近。<br/> <br/>它传递的正能量，应当是让人更爱阅读、更懂思考、更接纳差异，而不是用所谓“正确”的阅读方式，框定唯一的答案，消解阅读本身的自由与美好。<br/> <br/><span style="font-weight: bold;">二、别让“社交阅读”，绑架了你的热爱</span><br/> <br/>曾经见过太多人，把读书变成了一场“社交表演”：<br/> <br/>- 为了融入圈子，硬着头皮读自己不感兴趣的书；<br/>- 为了得到认可，刻意迎合别人的观点，不敢说出真实想法；<br/>- 为了显得“合群”，把大量时间花在讨论、打卡上，反而没好好读完几本书。<br/> <br/>他们害怕“不合群”，害怕“被孤立”，于是把读书变成了获取人脉、塑造人设的工具。可到头来，书没读透，心也累了——因为这份热爱，从一开始就被绑上了“社交”的枷锁。<br/> <br/>其实，阅读从来不需要“合群”。你可以一个人看书，一个人散步，一个人写日记，在独处的时光里和文字深度对话；你也可以偶尔走出去，遇见同频的人，分享彼此的感悟。<br/> <br/>但前提是：<span style="font-weight: bold;">你的阅读，永远要先取悦自己。</span><br/> <br/>如果一场读书会让你觉得压抑、疲惫，如果一群人的讨论让你觉得虚伪、刻意，那不如转身离开——真正的阅读，从来不需要勉强自己融入不属于自己的圈子。</p><p>正因如此，我前两年就把微信读书的排行榜给关了！</p><div><p style="display: inline-block;"><img data-ratio="2.210185185185185" data-type="jpeg" data-w="1080" style="height: auto !important;" src="https://wechat2rss.xlab.app/img-proxy/?k=1c84c898&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sytLY77RSv8T8zezAXF0qmCQsysASm2jRfAexCXfwSmUHedFAVkSuyMbXicO7TcmzdZ3kN8hWNGRkEcJPf3gLiaL7GZIWv4mumsg%2F640%3Fwx_fmt%3Djpeg"/></p></div><p> <br/><span style="font-weight: bold;">三、守住阅读的“私人与真诚”，才是对文字最好的尊重</span><br/> <br/>有人说，现在的读书氛围太“浮躁”了。大家忙着晒书单、晒笔记、晒感悟，却很少有人愿意沉下心来，好好读完一本书，好好和自己对话。<br/> <br/>或许我们都该慢下来：<br/> <br/>- 放下“必须分享”的执念，允许自己有“读不懂”“不喜欢”的时刻；<br/>- 放下“必须合群”的焦虑，允许自己独自享受阅读的宁静；<br/>- 放下“必须正确”的枷锁，允许自己有独一无二的感受与思考。<br/> <br/>读书，从来不是为了证明什么，也不是为了迎合谁。它是你和文字的私会，是你和自己的对话，是你在喧嚣世界里，为自己留的一方净土。<br/> <br/>别让读书，变成一场“正确”的表演。守住这份私人与真诚，让阅读回归本质，才是对文字最大的尊重！</p><p style="display: none;"><mp-style-type data-value="10000"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=aa154b46&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486065%26idx%3D1%26sn%3D916323140152e2c7b2ada48050b6a75a">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Mon, 16 Mar 2026 22:18:00 +0800</pubDate>
    </item>
    <item>
      <title>NDSS 2026 论文清单及摘要（上）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486060&amp;idx=1&amp;sn=2ed581b7ad4a96197103b393cdfea9a7</link>
      <description></description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span>2026-03-01 14:04</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=54d79b8e&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FtJDT9c8t2szRT9abIrMZdzNMhdPM5Qo3pMhzpGwiawthXIg0pTxapxu6eIOqVRmh9oibCU3hPdx7uYNW23KzPxmj6tZD9XnShkicSGgRAcRfLc%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p><span leaf="">PS：以前采集论文是通过写爬虫到调用LLM API完成内容生成的，</span><span leaf="">虽然有用到的大毛已做代码生成和翻译，但多少还是有一点点人工，怎么判断以及token付费。</span><span leaf="">但是现在很多自主Agent出来后（</span><span leaf="">都是claude code开的好头</span><span leaf="">），一切都变得更加简单和自动化。今天这篇文章，我是直接用GML agent模式，直接发送1条指令全自动搞定的，一切变得如此顺畅。</span></p><p style="text-align: center;" nodeleaf=""><img data-aistatus="1" class="rich_pages wxw-img" data-ratio="0.5120370370370371" data-s="300,640" data-type="png" data-w="1080" type="block" data-imgfileid="100002403" src="https://wechat2rss.xlab.app/img-proxy/?k=1f91c4fe&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_png%2FtJDT9c8t2sxsv5tydXEicpEKDfE89ibsib9IxvL5oQdoOnBR9Nxg6yibDAsmiaTqHxdXGbJ6dIEibibZ1XIjhU6pt4SPeffrkl0ETMOfCIXFYxOkn8%2F640%3Fwx_fmt%3Dpng%26from%3Dappmsg"/></p><p><span leaf="">当前你看到的这段内容，也是微信语音输入自动生成的，略有修改，同时支持五笔和拼音。现在大家都</span><span leaf="">用大模型，但是</span><span leaf="">整天打prompt也累，因此一直尝试想找到一种可替代敲字</span><span leaf="">的方式，感觉语音输入就是我要找的方法。试过很多工具，包括系统自带的语音输入、trae语音输入、闪电说、typeless等方法，语音识别都不够准，或者生成速度慢，或者翻墙账号登录 。相比之下，微信输入法速度和准确率相对更高，还是免费的，</span><span leaf="">不过有时也是会识别错误的</span><span leaf="">。之前看网上有人推荐另一款收费的语音输入wisper flow，我还没用过，还不太舍得为</span><span leaf="">打字付费。若有</span><span leaf="">更好的输入工具大家也可以推荐一下。</span></p><p nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="2" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><p cid="n2" mdtype="paragraph" style="box-sizing: border-box;text-align: left;" data-pm-slice="0 0 []"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">1、A Causal Perspective for Enhancing Jailbreak Attack and Defense</span></span></p><p cid="n3" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">揭示大型语言模型（LLMs）中&#34;越狱&#34;背后的机制对于提高其安全性和可靠性至关重要，然而这些机制仍 poorly understood。现有研究主要通过探测潜在表示来分析越狱提示，往往忽视了可解释提示特征与越狱发生之间的因果关系。在这项工作中，我们提出了 Causal Analyst，一个将 LLMs 集成到数据驱动因果发现中的框架，用于识别越狱的直接原因并利用它们进行攻击和防御。我们引入了一个包含七个 LLMs 上 35k 次越狱尝试的综合数据集，该数据集从 100 个攻击模板和 50 个有害查询中系统构建，并标注了 37 个精心设计的人类可读提示特征。通过联合训练基于 LLM 的提示编码和基于 GNN 的因果图学习，我们重建了从提示特征到越狱响应的因果路径。我们的分析显示，特定特征如&#34;积极角色&#34;和&#34;任务步骤数量&#34;是越狱的直接因果驱动因素。我们通过两个应用展示了这些见解的实际效用：（1）一个 Jailbreaking Enhancer，它针对识别出的因果特征，显著提高了在公共基准上的攻击成功率；（2）一个 Guardrail Advisor，它利用学习到的因果图从模糊查询中提取真实的恶意意图。包括基线比较和因果结构验证在内的广泛实验证实了我们因果分析的稳健性及其优于非因果方法的性能。我们的结果表明，从因果角度分析越狱特征是提高 LLM 可靠性的有效且可解释的方法。我们的代码可在 </span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/Master-PLC/Causal-Analyst" target="_blank">https://github.com/Master-PLC/Causal-Analyst</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf=""> 获取。</span></span></p><p cid="n4" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f797-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f797-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n6" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">2、A Deep Dive into Function Inlining and its Security Implications for ML-based Binary Analysis</span></span></p><p cid="n7" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">函数内联优化是现代编译器中广泛使用的一种转换技术，它根据需要将调用点替换为被调用函数的主体。虽然这种转换能提高性能，但它会显著改变机器指令和控制流图等静态特征，而这些特征对二进制分析至关重要。然而，尽管函数内联的影响广泛，其安全影响至今仍未得到充分探索。本文首次从基于机器学习的二进制分析角度对函数内联进行了全面研究。为此，我们剖析了LLVM成本模型中的内联决策流程，并探索了能够显著提高函数内联比例的编译器选项组合，我们将其称为极端内联。我们重点关注五种基于机器学习的安全二进制分析任务，使用20个独特模型系统评估它们在极端内联场景下的鲁棒性。大量实验揭示了几个重要发现：i) 函数内联尽管本意是良性转换，但可能间接或直接影响机器学习模型的行为，可能被用于规避判别式或生成式机器学习模型；ii) 依赖静态特征的机器学习模型对内联可能高度敏感；iii) 微妙的编译器设置可被利用来刻意规避二进制变体；iv) 内联比例在不同应用程序和构建配置中差异显著，这削弱了机器学习模型训练和评估中一致性的假设。</span></span></p><p cid="n8" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1872-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1872-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n10" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">3、A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems</span></span></p><p cid="n11" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于机器学习(ML)的恶意流量检测是一种有前景的安全范式。它能够识别各种高级攻击，优于基于规则的传统检测方法。然而，这些ML模型的鲁棒性在很大程度上尚未得到探索，从而使攻击者能够制作规避检测的对抗性流量样本。现有的规避攻击通常依赖于过于严格的条件(例如，加密协议、Tor或专用设置)，或需要针对目标的详细先验知识(例如，训练数据和模型参数)，这在现实世界的黑盒场景中是不切实际的。因此，硬标签黑盒规避攻击(即无需内部目标洞察即可适用于不同任务和协议)的可行性仍然是一个开放的挑战。为此，我们开发了NetMasquerade，它利用强化学习(RL)来操纵攻击流量，使其模仿良性流量并规避检测。具体而言，我们建立了一个名为Traffic-BERT的定制预训练模型，利用网络专用分词器和注意力机制来提取多样化的良性流量模式。随后，我们将Traffic-BERT集成到RL框架中，使NetMasquerade能够基于良性流量模式以最小修改有效操纵恶意数据包序列。实验结果表明，NetMasquerade能够在80种攻击场景下使暴力攻击和隐蔽攻击规避6种现有检测方法，攻击成功率超过96.65%。值得注意的是，它可以规避那些在经验上或可证明上能够抵抗现有规避攻击的方法。最后，NetMasquerade实现了低延迟的对抗流量生成，展示了其在实际场景中的实用性。</span></span></p><p cid="n12" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s916-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s916-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n14" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">4、A Unified Defense Framework Against Membership Inference in Federated Learning via Distillation and Contribution-Aware Aggregation</span></span></p><p cid="n15" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">联邦学习能够实现去中心化的模型训练而无需暴露原始数据，使其成为隐私保护机器学习的一种有前景的范式。然而，它仍然容易受到成员推理攻击(MIAs)的威胁，攻击者可以推断特定数据点是否包含在训练集中，这带来了严重的隐私风险并破坏了数据本地性。现有的针对MIAs的防御方法存在显著局限性：一些会导致性能大幅下降，而另一些则无法同时防御被动和主动攻击向量。为应对这些挑战，本文提出了一种统一的防御框架，能够在保护目标模型实用性的同时，同时减轻联邦学习中的被动和主动MIAs。首先，我们在教师模型训练过程中引入改进的熵正则化，以增强成员数据的不确定性，比标准正则化提供更强的推理攻击抵抗力。其次，我们利用条件变分自编码器(CVAE)生成类条件合成数据用于监督学生训练，这避免了敏感数据的直接暴露，并提供比无标记替代方案更好的实用性。最后，我们设计了一种感知贡献的聚合策略，根据实用性调整本地模型的影响力，减轻恶意客户端在模型聚合过程中的影响。在四个基准数据集上的实验结果表明，所提出的方法显著降低了各种成员推理攻击的成功率，优于现有的最先进防御方法。此外，它始终保持高模型精度，证明了其在实际联邦学习部署中的实用性。</span></span></p><p cid="n16" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s413-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s413-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n18" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">5、Abuse Resistant Traceability with Minimal Trust for Encrypted Messaging Systems</span></span></p><p cid="n19" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">加密消息系统为用户提供端到端安全，但阻碍了内容审核，使得打击网络滥用行为变得困难。可追溯性提供了一种有前景的解决方案，使平台能够识别消息的发起者或传播者，然而这种能力可能被滥用于对无辜消息进行大规模监控。为缓解这一风险，现有方法将可追溯性限制在由多个用户举报或处于预定义黑名单中的问题消息上。然而，这些解决方案要么过度信任特定实体（例如定义黑名单的方），要么依赖于同一平台运行的服务器之间不串通的不切实际假设。在本文中，我们提出了一种抗滥用的源追溯方案，将可追溯性分配给不同的现实世界实体。具体而言，我们形式化定义了其语法并证明了其安全属性。我们的方案实现了两个基本原则：最小信任原则，确保只要参与追溯的单一参与者是诚实的，即使其他所有参与者串通，追溯也不会被滥用；以及最小信息披露原则，防止参与者获取任何对追溯不必要的信息（例如通信方的身份）。我们使用Signal部署的技术实现了我们的方案，评估结果表明，它提供了与易受滥用的最新方案相当的性能。</span></span></p><p cid="n20" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f456-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f456-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n22" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">6、Accurate Identification of the Vulnerability-Introducing Commit based on Differential Analysis of Patching Patterns</span></span></p><p cid="n23" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">当在特定软件版本中发现漏洞时，追溯提交历史以准确识别引入该漏洞的首次提交（称为漏洞引入提交，VIC）至关重要。本文提出了一种基于漏洞修补模式差异分析的方法来准确识别VIC。首先，我们比较漏洞修补前后的两个文件，将补丁中与漏洞相关的语句分类为不同的修补模式，如编码错误、不适当的数据流、 misplaced语句和缺失的关键检查。然后，基于这些修补模式，我们从易受攻击的文件中提取漏洞关键语句序列，并将其与早期提交进行匹配，以确定引入提交。为了评估该方法的有效性，我们收集了一个包含6920个CVE和5,859,238个提交的数据集，数据来源于开源软件，包括Linux内核、MySQL和OpenSSL等。实验结果表明，该方法达到了94.94%的检测准确率和86.92%的召回率，显著优于现有方法。</span></span></p><p cid="n24" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s140-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s140-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n26" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">7、ACE: A Security Architecture for LLM-Integrated App Systems</span></span></p><p cid="n27" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">LLM集成应用系统通过系统LLM使用交错规划和执行阶段调用第三方应用来扩展大型语言模型（LLMs）的效用。这些系统引入了新的攻击向量，恶意应用可能导致规划或执行完整性受损、可用性中断或执行过程中隐私泄露。在本工作中，我们确定了影响LLM集成应用规划完整性以及执行完整性和可用性的新攻击，并在IsolateGPT（一种旨在缓解恶意应用攻击的最新解决方案）上展示了这些攻击。我们提出了Abstract-Concrete-Execute（ACE），一种新的LLM集成应用系统安全架构，为系统规划和执行提供安全保证。具体而言，ACE将规划分为两个阶段：首先仅使用可信信息创建抽象执行计划，然后使用已安装的系统应用将抽象计划映射为具体计划。我们通过结构化计划输出的静态分析验证了我们系统生成的计划满足用户指定的安全信息流约束。在执行过程中，ACE强制应用之间的数据和能力隔离，并确保执行按照可信的抽象计划进行。我们通过实验证明，ACE能够抵御InjecAgent和Agent Security Bench基准测试中的间接提示注入攻击以及我们新引入的攻击。我们还使用LangChain基准测试中的工具使用套件评估了ACE在实际环境中的实用性。我们的架构代表了使用系统安全原则强化基于LLM系统的重大进展。</span></span></p><p cid="n28" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s352-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s352-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n30" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">8、Achieving Interpretable DL-based Web Attack Detection through Malicious Payload Localization</span></span></p><p cid="n31" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Web攻击对Web应用构成重大威胁。虽然基于深度学习的系统已成为检测Web攻击的有前景解决方案，但其缺乏可解释性阻碍了在生产环境中的部署。现有的可解释性方法无法解释Web攻击，因为它们忽略了HTTP请求的结构信息。它们仅识别一些重要特征，这些特征安全操作人员难以理解，也无法指导他们采取有效应对措施。在本文中，我们提出了WebSpotter，实现了可解释的Web攻击检测，通过定位HTTP请求中的恶意载荷来增强现有的基于深度学习的检测方法。这一方法源于观察发现恶意载荷通常对检测模型的预测有显著影响。WebSpotter识别HTTP请求中每个字段的重要性，然后利用机器学习模型学习这种重要性与恶意载荷之间的相关性。此外，我们展示了WebSpotter如何通过自动生成WAF规则来协助安全操作人员缓解攻击。在两个公共数据集和我们新构建的数据集上进行的大量评估表明，WebSpotter显著优于现有方法，与基线相比，定位准确率至少提高了22%。我们还从CVE和实际Web应用中收集的真实世界攻击进行了评估，以说明WebSpotter在实际场景中的有效性。</span></span></p><p cid="n32" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1029-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1029-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n34" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">9、Achieving Zen: Combining Mathematical and Programmatic Deep Learning Model Representations for Attribution and Reuse</span></span></p><p cid="n35" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">先前的工作已经开发了能够从系统内存或程序二进制文件中提取通用格式的深度学习(DL)模型以进行安全分析的技术。不幸的是，这些技术忽略了模型重用和任何白盒分析技术所需的DL模型程序表示的恢复。针对这一问题，我们提出了一种新颖的恢复方法，并构建了原型系统ZEN，该系统能自动恢复DL模型的程序表示，补充了先前工作对数学表示的恢复。ZEN能够识别未知DL系统中相对于基础模型的新代码，并生成补丁，使得恢复的DL模型可以被重用。我们在21个最先进的DL模型上评估了ZEN，包括语言和视觉领域的模型，如Llama 3和YoloV10。ZEN能够以100%的准确度将自定义模型归因于其基础模型，实现了模型重用。</span></span></p><p cid="n36" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1628-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1628-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n38" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">10、Action Required: A Mixed-Methods Study of Security Practices in GitHub Actions</span></span></p><p cid="n39" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">GitHub Actions已成为主导的持续集成/持续交付(CI/CD)平台，但最近的SolarWinds和tj-actions/changed-files等供应链攻击凸显了此类系统中的关键安全漏洞。虽然GitHub提供了官方安全实践来缓解这些风险，但它们在现实世界中的实施程度仍不为人知。我们进行了一项混合方法研究，分析了338,812个公共仓库并对100多名开发者进行了调查，以了解GitHub Actions中的安全实践实施情况。我们的发现揭示了五个关键安全实践的实施率低得惊人，范围从0.6%到52.9%。我们确定了三个主要障碍：缺乏意识(高达71.6%的非采用者不了解这些实践)、对适用性的误解以及对运营成本的担忧。仓库特征，如组织所有权和最近的开发活动，与更好的安全实践实施显著相关。基于这些实证见解，我们得出了可行的建议，将干预策略与适当的自动化水平保持一致，改进通知设计以提高意识，加强平台和IDE级别的支持，并明确说明风险和适用性的文档。</span></span></p><p cid="n40" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f483-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f483-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n42" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">11、Actively Understanding the Dynamics and Risks of the Threat Intelligence Ecosystem</span></span></p><p cid="n43" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">尽管威胁情报（TI）生态系统已投入数十亿美元——这是一个由安全供应商和利他主义者组成的全球分布式网络，推动着关键网络安全运营——我们仍缺乏对其运作方式的理解，包括其动态和脆弱性。为填补这一空白，我们提出了一种新颖的测量框架，通过监控带有网络入侵指标（IoCs）水印的追踪二进制文件，来跟踪它们在生态系统中的传播。通过分析提交威胁情报的传播链的每个阶段（提交、提取、共享和阻断），我们发现一个生态系统，其中传播几乎总是导致威胁的阻断，但供应商选择性地共享他们提取的威胁情报，限制了生态系统的效用。此外，我们发现，试图遏制威胁的努力常常因&#34;瓶颈&#34;供应商延迟数小时至数天共享威胁情报而放缓。关键的是，我们确定了威胁情报供应链的多种威胁，其中一些目前已在野外被利用。供应商不必要的主动探测、对放置文件的浅层提取以及易于预测的沙箱环境指纹都威胁着生态系统的健康。为解决这些问题，我们为供应商和从业人员提供了可操作的改进威胁情报供应链安全的建议，包括已知滥用模式的检测特征。我们通过负责任的披露流程与供应商合作，了解了这些弱点背后的运营约束。最后，我们为积极测量威胁情报生态系统的研究人员提供了一套伦理最佳实践。</span></span></p><p cid="n44" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f102-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f102-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n46" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">12、ACTS: Attestations of Contents in TLS Sessions</span></span></p><p cid="n47" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Web3大规模应用的一个基本要求是使用户能够从其数据中受益，即使在已部署的系统内也是如此。这提出了一个重要的开放性问题：现有的、广泛采用的软件如何能够验证用户是否从TLS服务器检索了特定数据？最近，令人印象深刻的科学成果（例如DECO [CCS20]和Xie等人[USENIX24]的工作）和工业产品（TLSNotary）在上述具有挑战性的方向上取得了进展。然而，虽然这些方法很好地保持了TLS服务器不变，但检索到的数据随后被用于与验证者的计算中，而验证者需要运行一些先进的非标准化密码方案（例如ZK-SNARKs），这显然限制了所提出技术的大规模应用。</span></span></p><p cid="n49" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在本文中，我们基于先前的方法并依靠Fuchsbauer和Wolf [Eurocrypt24]提出的谓词盲签名这一新概念，通过提出ACTS（一种分布式架构）来绕过先前工作的局限性。ACTS仍然保持TLS服务器不变，同时允许用户证明其拥有从TLS服务器检索的数据，仅需验证者的软件能够检查标准签名即可。我们的贡献包括一个轮次最优的谓词盲签名协议，该协议生成标准的RSA-PSS签名。我们展示了如何将这一基本构件集成到DECO架构（及其后续版本）中，以证明从TLS服务器检索的数据。此外，我们已经优化了我们的构建，使其在商用硬件上对于公证人（即负责无意识认证TLS数据并保持数据保密性的参与者）实现的大而重要的策略类别是实用的。我们提供了一个实验评估，评估的场景是从TLS服务器下载的PDF文档并编码为AES-GCM密文。然后，用户将通过标准PADES签名获得一个经过认证的PDF，该签名由公证服务无意识地添加到PDF中，并附带一些元数据。生成的标准签名PDF文档可以使用现成的PDF阅读器透明验证。我们的实验验证表明，我们的架构适用于具体场景的实际部署。</span></span></p><p cid="n50" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1861-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1861-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n52" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">13、ADGFUZZ: Assignment Dependency-Guided Fuzzing for Robotic Vehicles</span></span></p><p cid="n53" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">机器人车辆（RV）在现代社会中扮演着越来越重要的角色，在商业和军事领域都有广泛应用。RV控制软件是RV系统的核心，它通过持续计算车辆的内部状态、传感器读数和外部输入来调整系统行为，从而保持正常运行。然而，RV软件中可配置参数、命令输入和环境感知数据的巨大组合空间给系统带来了显著的安全风险。现有的模糊测试技术在有效探索这一巨大输入空间的同时发现深层漏洞方面面临重大挑战。为应对这些挑战，我们提出了ADGFuzz，一种专门用于检测RV控制软件中赋值语句漏洞的新型模糊测试框架。ADGFuzz静态构建赋值依赖图（ADG）来捕获程序内的变量间依赖关系。然后，通过利用命名相似性将这些依赖关系传播到RV输入空间，从而产生一组称为匹配输入集（MIS）的定向输入。在此基础上，ADGFuzz在MIS上进行感知熵的模糊测试，从而提高漏洞发现的总体效率。在我们的评估中，ADGFuzz在三种RV类型中发现了87个独特漏洞，其中78个是先前未知的。所有发现的漏洞都已负责任地披露给开发人员，其中16个已被确认修复。</span></span></p><p cid="n54" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1014-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1014-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n56" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">14、AirSnitch: Demystifying and Breaking Client Isolation in Wi-Fi Networks</span></span></p><p cid="n57" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">为防止恶意Wi-Fi客户端攻击同一网络上的其他客户端，厂商引入了客户端隔离，这是一组阻止客户端之间直接通信的机制组合。然而，客户端隔离并非标准化功能，其安全保证尚不明确。在本文中，我们对Wi-Fi客户端隔离进行了结构化安全分析，发现了绕过此保护的新一类攻击。我们确定了这些弱点背后的几个根本原因。首先，保护广播帧的Wi-Fi密钥管理不当，可能被滥用以绕过客户端隔离。其次，隔离通常仅在MAC层或IP层执行，而非同时执行。第三，客户端身份在网络堆栈中的弱同步允许在网络层绕过Wi-Fi客户端隔离，从而能够拦截其他客户端以及内部后端设备的上行和下行流量。所有测试的路由器和网络都至少存在一种漏洞。更广泛地说，缺乏标准化导致各厂商实施的隔离措施不一致、临时且往往不完整。基于这些见解，我们设计并评估了端到端攻击，使现代Wi-Fi网络具备完整的中间人攻击能力。尽管客户端隔离有效缓解了诸如ARP欺骗等传统攻击，而ARP欺骗长期以来被认为是局域网中实现中间人定位的唯一通用方法，但我们的攻击提出了一种通用且实用的替代方案，即使在存在客户端隔离的情况下也能恢复这一能力。</span></span></p><p cid="n58" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1282-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1282-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n60" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">15、Aliens Among Us: Observing Private or Reserved IPs on the Public Internet</span></span></p><p cid="n61" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">伪造流量仍然是网络卫生的主要问题，因为它通过掩盖攻击来源并阻碍取证分析，使得分布式拒绝服务(DDoS)攻击成为可能。不良卫生状况的一个关键指标是公共互联网中存在&#34;虚假流量&#34;(Bogon traffic)——携带无效或不可路由源地址的数据包——这些数据包源于配置错误或过滤不足。尽管长期以来一直有源地址验证(SAV)的建议，如BCP 38和BCP 84，但虚假过滤的部署仍然不一致。在这项工作中，我们分析了CAIDA Ark平台八年间(2017-2024)的traceroute测量数据，并结合了RIPE RIS和RouteViews的历史BGP数据，以量化数据平面中虚假地址的普遍性和特征。我们观察到对最佳实践的广泛不遵守：在82.69%到97.83%的Ark观测点中，traceroute路径包含虚假IP地址，主要是RFC1918地址。总体而言，21.11%的traceroute包含RFC1918地址，较小比例涉及RFC6598(1.68%)和RFC3927(0.08%)。我们识别出超过15,500个传输虚假流量的自治系统(ASes)，但其中只有11.88%在超过一半的测量中这样做。与Spoofer项目和MANRS的交叉比对显示控制平面和数据平面保证之间存在显著差距：52.71%转发源自虚假数据包的ASes被分类为不可伪造，表明SAV部署不完整或无效。</span></span></p><p cid="n62" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1118-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1118-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n64" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">16、An LLM-Driven Fuzzing Framework for Detecting Logic Instruction Bugs in PLCs</span></span></p><p cid="n65" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可编程逻辑控制器(PLC)利用供应商提供的逻辑指令库（编译到设备固件中）来自动化工业操作。这些库可能包含安全漏洞，当通过物理控制例程、面向网络的服务或PLC运行时子系统利用时，可能导致权限违规、内存损坏或数据泄露。本文提出了LogicFuzz，这是首个专门针对PLC固件中逻辑指令设计的模糊测试框架。LogicFuzz构建了一个语义依赖图(SDG)，该图捕获了PLC代码中的操作语义和指令间依赖性。利用SDG和使能信号机制，LogicFuzz自动合成针对特定指令的种子程序，显著减少了手动工作量，并能够在真实PLC硬件上进行可控、可重置的模糊测试。为了发现依赖于控制流触发器（即调用模式）的缺陷，LogicFuzz对SDG进行变异以多样化指令调用上下文。为了暴露数据触发的故障，它在有效的语义约束下执行基于覆盖率的参数变异。此外，LogicFuzz集成了一个多源预言机，用于监控运行时日志、状态LED和通信状态，以在模糊测试期间检测指令级故障。我们在来自三大厂商的六款商用PLC上评估了LogicFuzz，发现了19个指令级漏洞，其中包括四个先前未知的漏洞。</span></span></p><p cid="n66" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1081-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1081-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n68" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">17、Analysis of the Security Design, Engineering, and Implementation of the SecureDNA System</span></span></p><p cid="n69" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们分析了SecureDNA系统在设计、工程和实现方面的安全方面。该系统使DNA合成器能够根据危险品数据库筛选订单请求。通过应用涉及分布式无意识伪随机函数的新密码学，该系统旨在保持订单请求和危险品数据库的机密性。我们从源代码（版本1.0.8）中部分识别了系统的详细操作，我们的分析检查了密钥管理、证书基础设施、身份验证和速率限制机制。我们还对相互认证、基本请求和豁免处理协议进行了首次形式化方法分析。在不破坏密码学的情况下，我们的主要发现是，SecureDNA的自定义相互认证协议SCEP仅实现了单向认证：危险品数据库和密钥服务器永远不知道它们与谁通信。这种结构性弱点违反了纵深防御原则，并使对手能够规避保护危险品数据库机密性的速率限制，前提是合成器连接到恶意或被破坏的密钥服务器或哈希数据库。我们指出了另一个违反纵深防御原则的结构性弱点：不足的密码绑定使系统无法检测TLS通道中来自危险品数据库的响应是否被修改。因此，如果合成器通过相同的TLS会话重新连接到数据库，对手可以重播和交换来自数据库的响应，而无需破坏TLS。尽管SecureDNA实现不允许此类重新连接，但避免潜在的结构性弱点将是更强的安全工程。我们确定了这些漏洞，并建议并验证了缓解措施，包括添加强绑定。我们的工作表明，一个安全的系统不仅需要健全的数学密码学，还需要形式化规范、健全的密钥管理、协议消息组件的适当绑定以及对工程和实现细节的谨慎关注。</span></span></p><p cid="n70" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1138-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1138-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n72" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">18、Anchors of Trust: A Usability Study on User Awareness, Consent, and Control in Cross-Device Authentication</span></span></p><p cid="n73" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">跨设备认证（XDAuth）已成为实现多设备无缝账户访问的关键机制。在此模式下，用户可以通过在另一个持有活跃会话或存储凭证的可信设备（认证设备）上完成认证来登录目标设备，从而提升用户体验。然而，认证设备与目标设备的分离引入了新的风险：物理和上下文的分离破坏了常规的认证流程，造成了信息不对称，并使用户难以评估认证请求的合法性。因此，用户可能会无意中批准恶意登录并导致账户被入侵，特别是在缺少关键上下文信息、明确确认机制或撤销功能的情况下。</span></span></p><p cid="n74" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">为解决这些风险，我们从以用户为中心的视角出发，基于三项基本用户权利（知情权、同意权和控制权）来保障XDAuth系统的安全性和可用性。我们通过研究27个采用三种典型XDAuth方案的主要服务，考察这些权利在实际应用中的支持情况。我们的发现令人担忧：超过一半的服务在认证过程中未提供任何关于目标设备的信息，并非所有服务都强制要求用户明确确认，且六个服务缺乏撤销可疑授权的途径。我们已负责任地向相关供应商披露了这些问题，其中多家供应商承认了问题并作出了积极回应。我们进一步对100名参与者进行了用户研究，发现绝大多数用户认为这些权利至关重要，并期望在XDAuth中得到保障。我们的研究揭示了当前实现与用户期望之间的明显差距，强调了需要加强对用户权利的支持，以开发更安全、以用户为中心的XDAuth系统。</span></span></p><p cid="n75" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f656-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f656-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n77" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">19、ANONYCALL: Enabling Native Private Calling in Mobile Networks</span></span></p><p cid="n78" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">移动网络运营商（MNOs）被曝出会泄露或出售用户的敏感信息，包括地理位置和通信历史记录。匿名移动用户认证方法，如文献{schmitt2021pretty}（USENIX Sec&#39;21）、文献{yu2023aaka}（NDSS&#39;24）、文献{alnashwan2024strong}（CCS&#39;24），使用户能够访问移动网络而不必暴露电话号码或订阅永久标识符（SUPI）等长期标识符。然而，身份透明度和位置感知的缺失在现实移动网络中实施匿名访问带来了重大挑战，尤其对于呼叫路由、使用量测量和计费等基本功能。为解决这些局限性，我们提出了ANONYCALL，一种隐私保护的呼叫管理架构，它支持匿名移动网络访问，同时实现两项基本功能：匿名被叫方发现和基于使用量的计费。ANONYCALL集成了一种带外认证机制，用于安全地共享临时呼叫标识符，实现无缝呼叫路由而不暴露永久用户信息。此外，它引入了一种匿名但可负责的余额凭证，能够实现准确计费并防止双重支付，同时保持移动用户匿名性。ANONYCALL完全兼容现有移动网络，引入的开销极小，呼叫建立时间增加不到200毫秒。通过智能手机和标准呼叫系统进行的评估证明了其实用性，为隐私保护且功能完备的移动通信提供了可行的解决方案。</span></span></p><p cid="n79" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1064-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1064-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n81" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">20、Anota: Identifying Business Logic Vulnerabilities via Annotation-Based Sanitization</span></span></p><p cid="n83" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">检测业务逻辑漏洞是软件安全中的一个关键挑战。这些漏洞源于应用程序设计或实现中的错误，允许攻击者触发非预期的应用程序行为。传统动态分析中的模糊测试净化器在发现与内存安全违规相关的漏洞方面表现出色，但 largely 无法检测业务逻辑漏洞，因为这些漏洞需要理解应用程序特定的语义上下文。最近尝试推断这种上下文的方法，由于依赖于启发式和非可移植的语言特性，本质上存在脆弱性和不完整性。由于业务逻辑漏洞构成了实践中最危险的软件弱点（CWE前40名中的27个）中的大多数，这是现有工具的一个令人担忧的盲点。在本文中，我们提出了一种名为ANOTA的新型人机交互净化框架来应对这一挑战。ANOTA引入了一个轻量级、用户友好的注释系统，使用户能够直接将其领域特定知识编码为轻量级注释，这些注释定义了应用程序的预期行为。然后，运行时执行监视器观察程序行为，将其与注释定义的策略进行比较，从而识别出表示漏洞的偏差。为了评估ANOTA的有效性，我们将ANOTA与最先进的模糊测试工具相结合，并与兼容相同目标的其他流行错误检测方法进行比较。结果表明，ANOTA+FUZZER在有效性方面优于这些方法。更具体地说，ANOTA+FUZZER能够成功重现43个已知漏洞，并在评估过程中发现了22个先前未知的漏洞（已分配17个CVE）。这些结果表明，ANOTA提供了一种实用且有效的方法，可以发现传统安全技术经常遗漏的复杂业务逻辑缺陷。</span></span></p><p cid="n84" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f938-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f938-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n86" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">21、Are your Sites Truly Isolated? Automatically Detecting Logic Bugs in Site Isolation Implementations</span></span></p><p cid="n87" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">站点隔离是现代浏览器的核心安全机制之一。通过将JavaScript即时编译器或HTML渲染等方面限制在沙盒进程中，网络浏览器显著减少了内存损坏错误的影响。此外，该机制还能防御Spectre等微架构攻击。使用站点隔离时，浏览器会将与特定站点相关的所有处理限制在其各自的沙盒进程中。与特权浏览器进程的所有通信都通过交换IPC消息完成。然而，这要求浏览器进程跟踪哪个渲染进程属于哪个站点，否则攻击者可能利用渲染器中的内存损坏问题，通过发送恶意IPC消息攻击其他站点。这反过来又可能允许攻击者泄露敏感数据（如cookies），甚至实现跨站脚本攻击（Universal Cross-Site Scripting）。</span></span></p><p cid="n88" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">本研究首次提出了在Firefox和Chrome中自动检测此类漏洞（称为站点隔离绕过漏洞）的方法。为此，我们提出了一种新的预言机制，通过标记进程级别的跨站点数据泄露来检测导致站点隔离绕过漏洞的语义错误。此外，我们还设计了一个模糊测试工具，模拟被攻陷的渲染进程，通过挂钩IPC通信，尝试利用浏览器进程作为受迷惑的代理。我们的研究在Chrome和Firefox中发现了四个安全漏洞：三个较轻微的漏洞会导致跨站点数据泄露，而第四个漏洞则允许完全控制目标站点。</span></span></p><p cid="n89" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f902-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f902-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n91" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">22、Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs</span></span></p><p cid="n92" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）已被集成到许多应用（如网络代理）中，以执行更复杂的任务。然而，由LLM驱动的应用程序容易受到间接提示注入（IPI）攻击的威胁，其中指令通过不可信的外部数据源被注入。本文提出了Rennervate，一个用于检测和预防IPI攻击的防御框架。Rennervate利用注意力特征在细粒度的令牌级别检测隐蔽注入，实现精确的净化，在保持LLM功能的同时中和IPI攻击。具体而言，令牌级检测器通过两步注意力池化机制实现，该机制聚合注意力头和响应令牌以进行IPI检测和净化。此外，我们建立了一个细粒度的IPI数据集FIPI，将开源以支持进一步研究。大量实验验证了Rennervate优于15种商业和学术IPI防御方法，在5个LLMs和6个数据集上实现了高精度。我们还证明了Rennervate可迁移到未见过的攻击，并能抵御自适应对手。</span></span></p><p cid="n93" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f394-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f394-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n95" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">23、Augmented Shuffle Differential Privacy Protocols for Large-Domain Categorical and Key-Value Data</span></span></p><p cid="n96" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">协议的新型增强型洗牌DP协议来填补这一空白。我们的FME协议使用哈希函数过滤掉不流行项目，然后准确计算流行项目的频率。为了在用户与洗牌者之间的一次交互轮次内完成此操作，我们的协议通过多重加密在系统内进行精心通信。我们还应用FME协议进行更高级的KV（键值）统计估计，并采用额外技术来减少偏差。对于分类数据和KV数据，我们证明了我们的协议提供了计算差分隐私，对上述两种攻击具有高度鲁棒性，同时保持了高精度和效率。通过与十二种现有协议的比较，我们展示了我们提案的有效性。</span></span></p><p cid="n97" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1124-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1124-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n99" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">24、Automated Code Annotation with LLMs for Establishing TEE Boundaries</span></span></p><p cid="n100" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代系统日益依赖可信执行环境（TEEs），如Intel SGX和ARM TrustZone，以安全地隔离敏感代码并减少可信计算基（TCB）。然而，识别应位于TEE中的精确代码区域，特别是涉及加密逻辑的代码区域，仍然具有挑战性，因为这需要深入的手动检查，且尚未得到自动化工具的支持。为解决这一开放性问题，我们提出了基于大型语言模型的代码标注逻辑（LLM-CAL），这是一种利用最新和先进的大型语言模型（LLMs）大规模自动化识别安全敏感代码区域的工具。我们的方法利用基础LLMs（Gemma-2B、CodeGemma-2B和LLaMA-7B），并通过使用新收集的包含4,000多个C源文件的手动标注数据集对这些模型进行了微调。我们将局部上下文特征、全局语义信息和结构元编码为紧凑的输入序列，引导模型捕捉代码中安全敏感性的微妙模式。微调过程基于量化LoRA——一种参数高效技术，在LLM架构中引入轻量级可训练适配器。为支持实际部署，我们开发了一个可扩展的数据预处理和推理流水线。LLM-CAL在识别敏感和非敏感代码方面达到了98.40%的F1分数和97.50%的召回率。这是首次尝试为启用TEE的平台自动化标注加密安全敏感代码，旨在最小化可信计算基（TCB）并优化TEE使用，以增强整体系统安全性。</span></span></p><p cid="n101" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s709-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s709-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n103" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">25、Automating Function-Level TARA for Automotive Full-Lifecycle Security</span></span></p><p cid="n104" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着现代汽车演变为智能互联系统，其日益增长的复杂性带来了重大的网络安全风险。因此，在强制性法规下，威胁分析与风险评估（TARA）已成为管理这些风险的关键手段。然而，现有的TARA自动化方法依赖于静态威胁库，限制了其在行业所需详细功能级分析中的应用。本文介绍了DefenseWeaver，这是首个利用组件特定细节和大语言模型（LLM）自动化功能级TARA的系统。DefenseWeaver从扩展的OpenXSAM++格式描述的系统配置中动态生成攻击树和风险评估，然后采用多智能体框架协调专门的LLM角色以实现更强大的分析能力。为进一步适应不断演变的威胁和多样化的标准，DefenseWeaver集成了低秩适应（LoRA）微调和基于专家策划的TARA报告的检索增强生成（RAG）。我们通过在四个汽车安全项目中的部署验证了DefenseWeaver的有效性，它识别出11条关键攻击路径，这些路径已通过渗透测试验证，并由相关汽车制造商和供应商进行了报告和修复。此外，DefenseWeaver展示了跨领域适应性，成功应用于无人机（UAV）和导航系统。与人类专家相比，在六种评估场景中，DefenseWeaver在手动攻击树生成方面表现更优。集成到UAES和小米等商业网络安全平台后，DefenseWeaver已生成超过8,200个攻击树。这些结果凸显了其显著减少处理时间的能力，以及其在各行业网络安全方面的可扩展性和变革性影响。</span></span></p><p cid="n105" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1408-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1408-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n107" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">26、BACnet or “BADnet”? On the (In)Security of Implicitly Reserved Fields in BACnet</span></span></p><p cid="n108" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">楼宇自动化系统（BAS）对于管理现代建筑中的供暖、通风、空调和制冷（HVAC&amp;R）以及照明和安全等基本功能至关重要。BACnet作为BAS广泛采用的开放标准，实现了异构设备之间的集成和互操作性。然而，传统的BACnet实现仍然容易受到各种安全威胁的攻击。虽然现有的模糊测试工具已被应用于BACnet，但其效率有限，主要原因是基于总线的通信介质速度慢且吞吐量低。为应对这些挑战，我们提出了BACsFuzz，一种行为驱动的模糊测试工具，旨在发现BACnet系统中的漏洞。与关注输入多样性和执行路径覆盖的传统模糊测试方法不同，BACsFuzz引入了令牌抢占辅助模糊测试技术，该技术利用BACnet的令牌传递机制提高模糊测试效率。令牌抢占辅助模糊测试技术被证明在发现由隐式保留字段滥用引起的漏洞方面非常有效。我们确定这是一个影响BACnet和KNX（另一种主要的BAS协议）的常见漏洞。值得注意的是，BACnet协会（ASHRAE）确认了协议级别的令牌抢占漏洞的存在，进一步验证了这一发现的重要性。我们在来自西门子、霍尼韦尔和江森自控等领先制造商的15个BACnet和5个KNX实现上评估了BACsFuzz。与最先进（SOTA）的方法相比，BACsFuzz将模糊测试吞吐量提高了272.49%至776.01%。总共发现了26个漏洞——18个在BACnet中，8个在KNX中——都与隐式保留字段相关。其中，24个漏洞已由制造商确认，9个已被分配CVE编号。</span></span></p><p cid="n109" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s794-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s794-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n111" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">27、Benchmarking and Understanding Safety Risks in AI Character Platforms</span></span></p><p cid="n113" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">AI角色平台允许用户与AI角色进行对话，是一个快速发展的应用领域。然而，其沉浸式和个性化的特点，加上技术漏洞，引发了重大的安全问题。尽管这些平台很受欢迎，但对其安全性的系统评估却明显缺失。为填补这一空白，我们进行了首个AI角色平台的大规模安全性研究，通过16个安全类别中的5000个基准问题对16个流行平台进行了评估。我们的研究结果揭示了一个关键的安全缺陷：AI角色平台的平均不安全响应率为65.1%，显著高于基线17.7%的平均水平。我们进一步发现，不同角色的安全性能差异显著，并与人口统计和性格等角色特征密切相关。利用这些见解，我们证明我们的机器学习模型能够以0.81的F1分数识别安全性较低的角色。这种预测能力对平台有益，能够促进更安全的交互机制、角色搜索/推荐和角色创建。总体而言，这些结果和发现为提升平台治理和内容审核以实现更安全的AI角色平台提供了宝贵的见解。</span></span></p><p cid="n114" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f575-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f575-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n116" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">28、Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services</span></span></p><p cid="n117" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在应用内应用生态系统中，超级应用为小程序开发者提供了访问各种敏感云服务的权限，如云数据库和云存储。这些服务使小程序开发者能够在超级应用服务器上高效地存储和管理小程序数据。为保护这些敏感资源，超级应用实施了身份管理机制，允许小程序开发者验证用户身份，确保只有授权和受信任的用户才能访问特定资源。然而，小程序开发者在资源管理实施中存在缺陷，可能导致敏感资源暴露给攻击者。在本文中，我们首次对应用内应用生态系统中的不安全云资源管理进行了系统性研究。我们设计并实现了一个名为ICREMiner的工具，该工具结合静态分析和动态探测技术，评估了在四个超级应用平台上访问应用内应用云服务的22,695个真实小程序的安全影响。研究结果显示，2,815个小程序（12.40%）受到不安全资源管理的影响，涉及8,062个不安全的云操作。我们发现一些知名企业的小程序也容易受到这些风险的威胁。此外，我们对该漏洞可能造成的重要安全危害进行了深入分析，例如允许攻击者窃取敏感用户信息和免费消费。作为回应，我们向超级应用平台和相应的小程序开发者进行了负责任的漏洞披露。我们还提供了几种缓解策略，帮助他们解决这些漏洞。</span></span></p><p cid="n118" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s194-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s194-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n120" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">29、Beyond Conventional Triggers: Auto-Contextualized Covert Triggers for Android Logic Bombs</span></span></p><p cid="n121" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">静态分析、模糊测试和基于学习的检测方面的最新进展显著提高了对触发型恶意软件的防御能力；然而，这些方法大多假设触发条件在语义上是明确的或与应用逻辑可区分的。在本文中，我们提出了SensorBomb，一种新颖的逻辑炸弹框架，它通过自动上下文化触发器和嵌入式传感器-执行器隐蔽信道利用了这一假设。SensorBomb不依赖于模糊或罕见的触发条件，而是构建与宿主应用的合法传感器使用、执行器行为和功能上下文紧密对齐的触发器，使其与良性行为无法区分。为此，SensorBomb自动分析宿主应用以选择兼容的传感器、执行器和敏感操作，构建隐蔽触发信道，并动态调整触发模式以逃避静态分析、模糊测试、传感器状态异常检测和用户怀疑。我们实现了三种此类触发器的代表性原型，并在不同设备和环境中进行了评估。结果表明，SensorBomb能够持续规避最先进的检测技术，实现高触发可靠性且无假阳性。对真实APK的大规模注入实验进一步证明，SensorBomb可以在不影响正常应用功能的情况下部署。这项工作揭示了移动恶意软件防御中一个关键且先前未被充分探索的攻击面，并呼吁开发更先进的检测机制。</span></span></p><p cid="n122" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f348-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f348-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n124" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">30、Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries</span></span></p><p cid="n125" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">LLM应用（即LLM应用程序）利用大语言模型的强大能力为用户提供定制化服务，革新了传统的应用开发模式。尽管日益普及的LLM驱动的应用为用户提供了前所未有的便利，但也带来了新的安全挑战。对于这样一个新兴的生态系统，安全界对LLM应用生态系统的理解尚不充分，特别是对应用自身能力边界的认识。在本文中，我们系统分析了新的开发范式，并定义了LLM应用能力空间的概念。我们还揭示了在现实场景中，由于能力边界模糊而可能产生的超越越狱攻击的新风险，即能力降级和能力升级。为评估这些风险的影响，我们设计并实现了一个LLM应用能力评估框架LLLMApp-Eval。首先，我们在4个平台上收集了应用元数据，并进行了跨平台生态系统分析。然后，我们对4个平台上的199个流行应用和6个开源大语言模型进行了风险评估。我们发现178个（89.45%）应用可能受到影响，这些应用能够执行来自15种以上场景的任务或具有恶意性。我们甚至在研究中发现了17个应用程序，它们直接执行恶意任务，而未应用任何对抗性重写。此外，我们的实验还揭示了提示设计质量与应用稳健性之间的正相关关系。我们发现精心设计的提示能增强安全性，而设计不佳的提示则可能助长滥用。我们希望我们的工作能够激励社区关注LLM应用的现实风险，促进更稳健的LLM应用生态系统的发展。</span></span></p><p cid="n126" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2941-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2941-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n128" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">31、Beyond Raw Bytes: Towards Large Malware Language Models</span></span></p><p cid="n129" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">恶意软件对关键计算基础设施构成日益增长的威胁，推动了对更先进的检测和分析方法的需求。尽管原始二进制恶意软件分类器显示出潜力，但其功能有限，难以应对长序列建模的挑战。与此同时，大型语言模型（LLMs）在自然语言处理领域的崛起展示了大规模、自监督模型在异构数据集上训练的力量，为众多下游任务提供了灵活的表示。这些模型成功的根源在于其训练数据的大小和质量、神经网络架构的表现力和可扩展性，以及其以自监督方式从未标记数据中学习的能力。在这项工作中，我们迈出了开发大型恶意软件语言模型（LMLMs）的第一步，这是LLMs在恶意软件领域的对应模型。我们解决了这一目标的核心方面，即关于数据、模型、预训练和微调的问题。通过使用语言建模目标预训练恶意软件分类模型，我们能够在各种实际的恶意软件分类任务上将下游性能平均提高1.1%，最高提高28.6%，这表明这些模型可以取代原始二进制恶意软件分类器。</span></span></p><p cid="n130" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s103-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s103-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n132" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">32、Beyond RTT: An Adversarially Robust Two-Tiered Approach For Residential Proxy Detection</span></span></p><p cid="n133" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">住宅IP代理网络已达到前所未有的规模，但它们通过将流量隐藏在合法的家庭地址背后，支持欺诈、网络抓取和复杂网络攻击等恶意活动，从而构成重大安全风险。现有的检测方法主要依赖于跨层往返时间（RTT）差异，但我们证明这些方法存在根本性缺陷：简单的流量调度攻击可以将检测召回率从99%降至仅8%，使得最先进的技术在面对基本对抗规避时变得不可靠。为解决这一关键漏洞，我们引入了新颖的流量分析和流关联特征，这些特征能够准确捕获网关和中继流量的特性，超越了易受攻击的基于时间的方法。我们进一步开发了CorrTransform，这是一种基于Transformer的深度学习架构，专为最大对抗弹性而设计。这实现了两种互补的检测策略：一种使用工程特征进行高效大规模检测的轻量级方法，以及一种在对抗环境中提供高保证的深度学习方法。我们通过对Bright Data的EarnApp进行为期15个月（900GB）涵盖超过110,000个代理连接的流量数据的综合分析，验证了我们的方法。我们的双层框架使ISP能够以&gt;98%的精确率/召回率识别代理设备，在正常条件下以99%的精确率/召回率分类单个连接，同时在对包括调度、填充和数据包重塑在内的复杂攻击保持&gt;92%的F1分数，而现有方法在这些攻击面前完全失效。对于内容提供商，我们的方法在区分直接流量与代理流量时实现了接近完美的召回率，同时假阳性率&lt;0.2%。这项工作将代理检测从易受攻击的基于时间的方法转变为具有弹性的架构指纹识别，为应对日益增长的恶意住宅代理使用威胁提供了可立即部署的工具。</span></span></p><p cid="n134" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2086-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2086-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n136" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">33、BINALIGNER: Aligning Binary Code for Cross-Compilation Environment Diffing</span></span></p><p cid="n138" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">二进制差异比对旨在对齐两个二进制文件中对应相同源代码片段的控制流图部分，以用于软件安全分析，如漏洞和抄袭检测任务。先前的工作在跨编译环境场景中效果有限且支持不够灵活。主要原因是它们基于基本块的相似性比较进行匹配。在我们的工作中，我们提出了一种新的二进制级别差异比对方法BINALIGNER，以缓解上述局限性。为了减少对应相同源代码片段的错误匹配和漏匹配的可能性，我们提出了条件松弛策略来寻找候选子图对。为了支持跨编译环境场景中更灵活的二进制差异比对，我们使用指令无关的基本块特征进行子图嵌入生成。我们实现了BINALIGNER，并在四种跨编译环境场景（即跨版本、跨编译器、跨优化级别和跨架构）中进行了实验，以评估其有效性和对不同场景的支持能力。实验结果表明，在大多数场景中，BINALIGNER显著优于最先进的方法。特别是在跨架构场景和跨编译环境场景的多种组合中，BINALIGNER的F1分数平均比基线方法高出65%。使用真实世界漏洞和补丁的两个案例研究进一步证明了BINALIGNER的实用性。</span></span></p><p cid="n139" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s649-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s649-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n141" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">34、Bit of a Close Talker: A Practical Guide to Serverless Cloud Co-Location Attacks</span></span></p><p cid="n142" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">无服务器计算通过为用户提供一种高效、经济的应用开发和部署方式而无需管理基础设施细节，从而彻底改变了云计算。然而，无服务器云用户仍然容易受到各种类型的攻击，包括微架构侧信道攻击。这些攻击通常依赖于受害者和攻击者实例的物理共存，攻击者需要利用云调度器来实现与受害者的共存。因此，研究无服务器云调度器的漏洞并评估不同无服务器调度算法的安全性至关重要。本研究解决了理解和构建无服务器云中共存攻击的空白问题。我们提出了一个全面的方法论，用于发现无服务器调度算法中的可利用特征，并制定通过正常用户界面构建共存攻击的策略。在我们的实验中，我们成功揭示了可利用的漏洞，并在流行的开源基础设施和微软Azure函数上实现了实例共存。我们还提出了一种缓解策略——双调度器（Double-Dip scheduler），以防御无服务器云中的共存攻击。我们的工作强调了当前云调度器中安全增强的关键领域，为加强无服务器计算环境抵御潜在的共存攻击提供了见解。</span></span></p><p cid="n143" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1376-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1376-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n145" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">35、BKPIR: Keyword PIR for Private Boolean Retrieval</span></span></p><p cid="n146" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">私有信息检索（Keyword PIR）使用户能够从数据库中检索与特定关键词相关的数据，同时保持查询的私密性。然而，现有的关键词PIR方案难以支持布尔检索模型，而该模型是实际应用中需要术语逻辑组合所必需的。本文提出了一种新颖的关键词PIR方案，利用了同态等值运算的进展。它支持在具有多对多关键词-值映射的数据库上进行隐私保护检索，同时支持布尔运算符以实现表达性搜索逻辑。重要的是，这种扩展保留了经典PIR的核心安全保证。据我们所知，这是首次将关键词PIR与布尔检索模型相结合的工作。实验评估表明，我们的方案实现了与多对多关键词-值数据库中值总数成比例的通信成本降低，同时获得了与值数量线性扩展的聚合查询处理性能提升。这些改进增强了其在隐私保护网络搜索和专利检索等实际应用中的可行性。</span></span></p><p cid="n147" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s536-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s536-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n149" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">36、Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks</span></span></p><p cid="n150" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）仍然容易受到越狱攻击的威胁，这些攻击利用对抗性提示来规避安全措施。当前的安全微调方法面临两个关键限制。首先，它们往往难以在安全性和实用性之间取得平衡，更强的安全措施往往会过度拒绝无害的用户请求。其次，它们经常忽略隐藏在看似良性任务中的恶意意图，使模型容易受到攻击。我们的工作确定了这些问题的根本原因：在响应生成过程中，LLM区分有害输出和安全输出的能力会减弱。实验证据证实了这一点，揭示出安全响应和有害响应的隐藏状态之间的可分性随着生成过程的推进而降低。这种减弱的辨别力迫使模型在生成过程的更早阶段做出合规性判断，限制了它们识别正在形成的恶意意图的能力，并导致了上述两种失败。为了缓解这一漏洞，我们引入了DEEPALIGN——一种增强LLM安全性的内在防御框架。通过在响应生成的中点应用对比隐藏状态引导，DEEPALIGN放大了有害和良性隐藏状态之间的分离，使生成过程中能够持续进行内在毒性检测和干预。此外，它有助于对有害查询提供上下文适当的安全响应，从而扩展安全响应的可行空间。评估结果表明了DEEPALIGN的有效性。在跨越不同架构和规模的多样化LLM中，它将九种不同越狱攻击的成功率降低到接近零或最低水平。重要的是，它在保持模型能力的同时减少了过度拒绝。配备DEEPALIGN的模型在拒绝具有挑战性的良性查询时，错误率降低了高达3.5%，并且标准任务性能下降不到1%。这标志着在安全-效用帕累托前沿方面取得了重大进展。</span></span></p><p cid="n151" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f4-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f4-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n153" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">37、BLERP: BLE Re-Pairing Attacks and Defenses</span></span></p><p cid="n155" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">蓝牙低功耗(BLE)是一种无处不在的无线技术，被数十亿设备用于交换敏感数据。根据蓝牙核心规范v6.1的定义，BLE的安全性依赖于两个主要协议：配对协议，用于建立长期密钥；以及会话建立协议，用于使用新的会话密钥加密通信。尽管标准允许已配对设备重新配对以协商新的安全级别，但这种机制的安全影响仍未被探索，尽管存在设备伪装和中间人(MitM)攻击的相关风险。我们分析了标准v6.1中定义的BLE重新配对机制，并确定了六个设计漏洞，其中包括四个新发现的漏洞，如未经验证的重新配对和安全级别降级。这些漏洞是设计缺陷，影响任何使用配对的符合标准的BLE设备，无论其蓝牙版本或安全级别如何。我们还提出了四种利用这些漏洞的新型重新配对攻击，我们称之为BLERP。这些攻击能够以最小或无需用户交互(一键或零点击)的方式实现设备伪装和中间人攻击。我们的攻击是首个针对BLE重新配对的攻击，利用了BLE配对与会话建立之间的相互作用，并滥用了SMP安全请求消息。我们开发了一个新型工具包，实现了我们的攻击并支持BLE配对的测试，包括端到端的中间人攻击。重现该工具包仅需低成本硬件(nRF52)和开源软件(Mynewt、NimBLE和Scapy)。我们的大规模评估展示了攻击对22个目标的影响，包括15个BLE主机、12个BLE控制器、高达5.4版本的蓝牙以及最安全的配置(SC、SCO和认证配对)。在我们的实验中，我们还发现了影响Apple、Android和NimBLE BLE栈的实现重新配对漏洞。我们实施并评估了两种互补的缓解措施：一种向后兼容的重新配对逻辑加固方案，供应商可立即部署；以及一种认证重新配对协议，从设计上解决了这些攻击。我们通过实证验证了加固重新配对的有效性，并使用ProVerif形式化建模和验证了认证重新配对。</span></span></p><p cid="n156" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f121-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f121-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n158" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">38、Breaking Isolation: A New Perspective on Hypervisor Exploitation via Cross-Domain Attacks</span></span></p><p cid="n159" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">虚拟机监控程序面临关键内存安全漏洞的威胁，其中指针损坏是最普遍和最严重的形式之一。现有的利用框架依赖于识别宿主机中的高度受限结构并准确确定其运行时地址，但在虚拟机监控程序环境中，此类结构稀少且被地址空间布局随机化(ASLR)进一步混淆，因此这种方法无效。我们观察到现代虚拟化环境存在弱内存隔离问题——客户机内存完全由攻击者控制，但可从宿主机访问，这为利用提供了可靠的原始基础。基于这一观察，我们首次对跨域攻击(CDA)进行了系统性的特征描述和分类，这是一类通过重用客户机内存实现能力提升的利用技术。为自动化这一过程，我们开发了一个系统，用于识别跨域小工具，将其与损坏的指针匹配，合成触发输入，并组装完整的利用链。我们在QEMU和VirtualBox的15个真实世界漏洞上的评估表明，CDA具有广泛的适用性和有效性。</span></span></p><p cid="n160" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f376-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f376-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n162" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">39、Breaking the Bulkhead: Demystifying Cross-Namespace Reference Vulnerabilities in Kubernetes Operators</span></span></p><p cid="n163" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Kubernetes Operator是一种旨在管理Kubernetes集群内应用生命周期的自动化工具，它扩展了Kubernetes的功能，并减轻了人类工程师的操作负担。虽然Operator显著简化了DevOps工作流程，但也引入了新的安全风险。特别是，Kubernetes强制执行命名空间隔离以分离工作负载并限制用户访问，确保用户只能与其授权命名空间内的资源交互。然而，Kubernetes Operator通常需要提升的权限，并且可能与多个命名空间中的资源交互。这引入了一类新的漏洞——跨命名空间引用漏洞。其根本原因在于资源声明的范围与Operator逻辑实现范围之间的不匹配，导致Kubernetes无法正确隔离命名空间。利用此类漏洞，具有单个授权命名空间有限权限的攻击者可能利用Operator执行影响其他未授权命名空间的操作，导致权限提升及其他进一步影响。据我们所知，本文是首个系统性研究Kubernetes Operator攻击的论文。我们提出了跨命名空间引用漏洞及其两种攻击策略，展示了攻击者如何绕过命名空间隔离。通过大规模测量，我们发现野外环境中超过14%的Operator可能存在漏洞。我们的发现已报告给相关开发者，截至投稿时已获得8项确认和7个CVE编号，影响了包括Kubernetes发明者谷歌和Operator发明者红帽在内的供应商，这凸显了增强Kubernetes Operator安全实践的迫切需求。为缓解此问题，我们开源了静态分析工具套件，并提出了具体的缓解措施以造福整个生态系统。</span></span></p><p cid="n164" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f761-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f761-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n166" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">40、Breaking the Generative Steganography Trilemma: ANStega for Optimal Capacity, Efficiency, and Security</span></span></p><p cid="n167" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">生成式隐写术在隐蔽通信方面展现出巨大潜力，然而现有方法常受限于容量、效率和安全性三者的权衡困境。基于霍夫曼编码（HC）的方法效率低下且安全性不足，而基于算术编码（AC）的方法虽能实现最优容量，但也存在安全风险。尽管近期已有可证明安全的方法解决了安全问题，但往往以增加嵌入复杂度或降低容量为代价——无法达到基于AC方法的高容量水平。为解决这一三重困境，我们将非对称数值系统（ANS）应用于隐写术。我们的核心洞见是重新利用ANS状态机，将其解码函数用于嵌入，编码函数用于提取。为将这一概念转化为实用系统，我们引入了几项关键创新。首先，我们采用流式架构结合状态重归一化，以实现任意长度消息的稳定嵌入。其次，我们采用直接浮点运算，避免高概率到频率的转换，从而降低复杂度和精度损失。更重要的是，我们引入了一种创新的密码学掩码机制，确保采样过程由密码学安全的伪随机数生成器驱动，从而实现可证明的安全性。最后，通过将核心计算优化为高效的位移操作，ANStega实现了卓越的嵌入和提取速度。实验结果验证了ANStega同时实现了最优嵌入容量、最优效率（O(1)嵌入复杂度）和最优安全性，成功解决了生成式隐写术中长期存在的三重困境。</span></span></p><p cid="n168" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f605-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f605-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n170" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">41、BSFuzzer: Context-Aware Semantic Fuzzing for BLE Logic Flaw Detection</span></span></p><p cid="n171" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">蓝牙低功耗（BLE）已成为现代互联设备的基础通信标准。然而，其复杂设计引入了微妙的逻辑缺陷，如字段误解或无效状态转换，这些缺陷可能导致身份验证绕过、未经授权的控制或拒绝服务（DoS）攻击。这些问题常常逃避传统的模糊测试和形式化分析。为解决这一差距，我们提出了BSFuzzer，一种基于蓝牙核心规范指导的、黑盒的、上下文感知的语义模糊测试框架。BSFuzzer利用大型语言模型（LLM）代理来语义解析蓝牙规范，从文本、图表和上下文中提取状态机和数据包语义。然后生成两种类型的变异：协议规则的字段级违规和关键转换的状态级破坏。这些变异被组合成结构化测试序列并在目标设备上执行。LLM代理进一步用于验证响应是否符合预期行为，从而能够检测传统模糊测试器无法触及的微妙逻辑缺陷。我们在19个真实的BLE设备上评估了BSFuzzer，包括9个系统级芯片（SoC）模块和10部智能手机。它发现了36个安全问题，其中包括34个先前未知的漏洞，其中9个已获得CVE标识符。两个关键漏洞通过漏洞赏金计划被一家主要供应商认可。实验结果表明，BSFuzzer在基于LLM的规范分析（高达97%）和响应验证（高达85.8%）方面均达到高准确率，证明了其在语义提取和提升模糊测试性能方面的有效性。与四种最先进的BLE漏洞检测工具相比，BSFuzzer实现了9.34%更高的代码覆盖率，并暴露了更广泛的漏洞类别，证明了其在发现BLE协议实现中深层解释不一致方面的有效性。</span></span></p><p cid="n172" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f94-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f94-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n174" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">42、Bullseye: Detecting Prototype Pollution in NPM Packages with Proof of Concept Exploits</span></span></p><p cid="n175" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">原型污染是JavaScript中的一个关键安全漏洞，特别是在Node.js包和应用程序中，攻击者可以操纵全局对象原型并向所有继承自该原型的对象注入恶意属性。最先进的静态和动态方法在检测此漏洞方面面临显著限制，无论是在准确性还是效率方面。静态方法难以识别不可利用的漏洞（例如，由于缺少带有预防机制的代码上下文），导致高误报率，同时还存在可扩展性问题。动态方法由于能够访问运行时信息，因此误报率较低；然而，由于代码可达性低（例如，由于使用了不适当的参数类型/值），其漏报率可能很高。在本文中，我们提出了Bullseye，一个全自动化动态分析框架，可对Node.js包中的原型污染漏洞提供经过验证且可扩展的分析。Bullseye的创新方法结合了广泛的入口点覆盖、上下文感知的漏洞生成和双运行时验证预言机。我们使用包测试套件中开发者提供的输入，以及从先前工作中提取的原型污染相关漏洞利用输入。然后，我们使用相关的漏洞利用输入候选执行每个入口点，并观察运行时以检测原型污染的迹象。我们在不到8小时内分析了44,513个高流行度的Node.js包（每周下载量超过10,000次）以及5,879个每周下载量较低的包。我们在290个包中检测到了零日原型污染漏洞，且没有误报。我们已负责任地向各包维护者披露了所有发现，并附带了概念验证漏洞利用代码。截至2025年7月22日，我们总共被分配了149个CVE；其中，66个已公开，25个被评为严重级别，34个被评为高危级别。</span></span></p><p cid="n176" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s211-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s211-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n178" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">43、BunnyFinder: Finding Incentive Flaws for Ethereum Consensus</span></span></p><p cid="n179" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">以太坊作为领先的区块链平台，依赖激励机制来提高其稳定性。最近，针对这些激励机制已出现多种攻击手段，例如所谓的重组攻击，这种攻击会导致诚验证者提出的区块被丢弃。在重组攻击中，诚验证者获得的奖励低于其应得的份额。然而，发现这些攻击严重依赖专业知识，且可能需要大量人工工作。我们提出了proto，一个只需少量人工工作即可发现以太坊激励缺陷的框架。proto受故障注入启发，这是一种在软件测试中常用的发现实现漏洞的技术。与发现实现漏洞不同，我们的目标是发现设计缺陷。我们的主要技术贡献包括一个精心设计的&#34;策略生成器&#34;，可生成大量攻击实例；一个自动工作流程，用于发起攻击并分析结果；以及一个集成了强化学习的工作流程，用于微调攻击参数并识别最具盈利能力的攻击。我们使用该框架模拟了总计7,991个攻击实例，并得出以下结果：首先，我们的框架重现了五种先前通过人工方式发现的已知激励攻击；其次，我们发现了三种可归类为激励缺陷的新攻击；最后且令人惊讶的是，我们的一个实验还发现了两个实现漏洞。</span></span></p><p cid="n180" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s281-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s281-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n182" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">44、Cache Me, Catch You: Cache Related Security Threats in LLM Serving Frameworks</span></span></p><p cid="n183" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）正在迅速重塑数字交互方式。其性能和效率高度依赖于先进的缓存机制，如前缀缓存和语义缓存。然而，这些机制引入了新的攻击面。与以往专注于训练阶段LLMs投毒攻击的研究不同，本文首次对LLM推理阶段出现的缓存相关安全风险进行了全面研究。我们对主流LLM服务框架中的缓存实现进行了系统性研究，随后确定了六种新型攻击向量，分为两类：（1）面向用户的欺诈攻击，通过前缀缓存碰撞和语义模糊投毒来操纵缓存条目，向用户传递恶意内容；（2）系统完整性攻击，利用缓存漏洞绕过安全检查，例如使用分块或多模态碰撞来规避内容审核。我们在领先的开源框架上验证了这些攻击向量，并评估了其影响和成本。此外，我们提出了五种多层防御策略并评估了其有效性。我们向受影响的供应商（包括vLLM、SGLang、GPTCache、AIBrix、rtp-llm和LMDeploy）负责任地披露了我们的发现。所有供应商都已确认这些漏洞，值得注意的是，vLLM、GPTCache和AIBrix已采纳我们提出的缓解方法并修复了其漏洞。我们的研究结果强调了在快速扩展的LLM生态系统中保护缓存基础设施的重要性。</span></span></p><p cid="n184" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2812-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2812-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n186" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">45、Cascading and Proxy Membership Inference Attacks</span></span></p><p cid="n187" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">成员推理攻击（MIA）通过确定特定查询实例是否包含在数据集中，来评估训练好的机器学习模型对其训练数据的揭示程度。根据攻击者是否被允许在成员查询上训练影子模型，我们将现有的MIA分为自适应或非自适应两类。在自适应设置中，攻击者在访问查询实例后可以训练影子模型，我们强调了利用实例间成员依赖关系的重要性，并提出了一种称为级联成员推理攻击（CMIA）的攻击无关框架，该框架通过条件影子训练整合成员依赖关系，以提高成员推理性能。在非自适应设置中，攻击者被限制在获取成员查询前训练影子模型，我们引入了代理成员推理攻击（PMIA）。PMIA采用代理选择策略，识别与查询实例行为相似的样本，并利用它们在影子模型中的行为进行成员后验概率测试以执行成员推理。我们为这两种攻击提供了理论分析，大量实验结果表明，在两种设置下，CMIA和PMIA都显著优于现有的MIA，特别是在低假阳性区域，这对于评估隐私风险至关重要。</span></span></p><p cid="n188" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s661-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s661-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n190" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">46、CAT: Can Trust be Predicted with Context-Awareness in Dynamic Heterogeneous Networks?</span></span></p><p cid="n191" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">信任预测为决策制定、风险缓解和系统安全增强提供了有价值的支持。最近，图神经网络（GNN）已成为一种有前景的信任预测方法，因为它能够学习能够捕捉网络内复杂信任关系的 expressive 节点表示。然而，当前基于GNN的信任预测模型面临几个局限性：（i）大多数模型无法捕捉信任的动态性，导致推理结果存疑。（ii）它们很少考虑现实网络的异构性，导致丰富语义的丢失。（iii）它们都不支持上下文感知性，这是信任的基本属性，使得预测结果变得粗糙。为此，我们提出了CAT，这是第一个支持信任动态性并能准确表示现实世界异构性的基于GNN的上下文感知信任预测模型。CAT包含图构建层、嵌入层、异构注意力层和预测层。它使用连续时间表示处理动态图，并通过时间编码函数捕捉时间信息。为了建模图的异构性并利用语义信息，CAT采用双重注意力机制，识别不同节点类型以及每种类型内节点的重要性。为了实现上下文感知，我们引入了元路径的新概念来提取上下文特征。通过构建上下文嵌入和集成上下文感知聚合器，CAT可以预测上下文感知信任和整体信任。在三个真实数据集上的广泛实验表明，CAT在信任预测方面优于五组基线方法，同时展现出对大规模图的强大可扩展性以及对信任导向和GNN导向攻击的鲁棒性。</span></span></p><p cid="n192" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2171-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2171-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n194" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">47、CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding</span></span></p><p cid="n195" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">机器学习中的后门攻击因其能够秘密破坏模型而引起了广泛关注，但大多数研究都集中在图像等 homogeneous 数据上。在这项工作中，我们提出了一种针对表格数据的新型后门攻击，由于同时存在数值和分类特征，这种攻击尤其具有挑战性。我们的核心思想是一种新颖的将分类值转换为浮点表示的技术。与传统方法如独热编码或序数编码相比，这种方法保留了足够的信息以保持干净模型的准确性。通过这种方法，我们创建了一种基于梯度的通用扰动，适用于所有特征，包括分类特征。我们在五个数据集和四种流行模型上评估了我们的方法。结果表明，在白盒和黑盒设置（包括 Vertex AI 等实际应用）中，攻击成功率高达100%，揭示了表格数据存在严重漏洞。我们的方法在性能上超越了先前的工作（如 Tabdoor），同时能够躲避最先进的防御机制。我们针对频谱签名、神经网络净化、Beatrix 和精细剪枝等防御方法评估了我们的攻击，所有这些方法都无法成功防御。我们还验证了我们的攻击能够成功绕过流行的异常检测机制。</span></span></p><p cid="n196" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1469-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1469-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n198" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">48、Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models</span></span></p><p cid="n1106" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">低秩适应（LoRA）已成为微调大型语言模型（LLMs）的有效方法，并在开源社区中得到广泛应用。然而，通过Hugging Face等平台分发LoRA适配器会带来新的安全漏洞：恶意适配器可以轻易传播并规避传统监督机制。尽管存在这些风险，针对基于LoRA的微调的后门攻击研究仍然相对不足。现有的后门攻击策略并不适用于此场景，因为它们通常依赖于无法获取的训练数据，未能考虑LoRA特有的结构特性，或遭受高误触发率（FTR），从而损害了其隐蔽性。为应对这些挑战，我们提出了因果引导的去毒化后门攻击（CBA），这是一种专为开源权重LoRA模型设计的新型后门攻击框架。CBA无需访问原始训练数据，并通过两项关键创新实现高度隐蔽性：（1）一种覆盖引导的数据生成流程，通过行为探索合成任务对齐的输入；（2）一种因果引导的去毒化策略，通过保留任务关键神经元合并中毒和干净的适配器。与先前方法不同，CBA能够基于因果影响进行权重分配，实现训练后的攻击强度控制，无需重复重新训练。在六个LoRA模型上的评估表明，CBA实现了高攻击成功率，同时将FTR比基线方法降低50-70%。此外，它对最先进的后门防御表现出更强的抵抗力，凸显了其隐蔽性和鲁棒性。</span></span></p><p cid="n200" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f168-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f168-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n202" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">49、Cease at the Ultimate Goodness: Towards Efficient Website Fingerprinting Defense via Iterative Mutual Information Minimization</span></span></p><p cid="n203" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">针对日益增长的网络隐私威胁，Tor网络通过去中心化、加密的基础设施路由流量，提供了针对监控的重要保护。然而，网站指纹攻击（WFA）对Tor的匿名性构成了严峻挑战。本文介绍了FRUGAL，一种利用网站流量与标签之间的互信息（MI）减少作为优化目标的流量混淆方法，为网站指纹防御（WFD）研究提供了新的视角。FRUGAL通过在最能累积减少互信息的位置 strategically 注入虚拟数据包，与最先进的（SOTA）防御机制相比取得了显著性能。它能在有效降低各种攻击模型下的攻击成功率（ASR）的同时，保持最小的带宽开销（BWO），并减轻对抗训练的影响。大量实验验证了FRUGAL在包括封闭世界、开放世界和真实世界模拟环境在内的各种场景中的有效性。例如，在封闭世界环境中，FRUGAL将DF模型的ASR降低至2.68%，带宽开销为30%，显著优于之前的SOTA防御方法，如Palette（11.54%的ASR，87%的BWO）。当FRUGAL的BWO增加到可比的80%水平时，ASR进一步降至1%以下，显示出显著的鲁棒性，即使在对抗训练后仍保持在9.42%，而Palette则为20.27%。这项研究不仅为WFD研究提供了新视角，还将FRUGAL确立为对抗WFA的强大通用防御框架。</span></span></p><p cid="n204" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f786-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f786-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n206" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">50、CELLSHIFT: RTT-Aware Trace Transduction for Real-World Website Fingerprinting</span></span></p><p cid="n207" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">网站指纹识别是一种隐私攻击，攻击者通过机器学习预测用户通过Tor网络访问的网站。近期研究提出使用Tor出口中继可测量的用户自然交互的&#34;真实&#34;模式或轨迹来评估WF攻击，但这些轨迹并不能准确反映入口侧WF攻击者所观察到的模式。在本文中，我们提出了将出口轨迹转换为入口轨迹的新方法，以便更准确地估计WF对实际Tor用户构成的风险。我们的方法利用轨迹时间戳和元数据提取多次往返时间估计，并使用它们将&#34;转换&#34;轨迹到目标观察点的视角。通过广泛评估，我们证明我们的方法在多个合成和真实数据集上均优于现有技术，且效率显著提高；它们使研究人员能够更准确地代表入口侧WF攻击者面临的现实挑战，并生成增强数据集，使攻击者能够提升现有WF攻击的性能。</span></span></p><p cid="n208" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1004-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1004-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n210" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">51、CHAMELEOSCAN: Demystifying and Detecting iOS Chameleon Apps via LLM-Powered UI Exploration</span></span></p><p cid="n211" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">变色龙应用在提交时展示合法功能以规避iOS应用商店审核，然后在安装后转变为非法版本。尽管这类应用普遍存在，但其底层转换方法和开发者-用户合谋机制仍鲜为人知。现有检测方法受限于静态分析或元数据依赖，对混合实现、新型变种或元数据稀缺实例无效。为解决这些局限，我们通过隐蔽渠道收集了500个iOS变色龙应用，构建了一个精心策划的数据集，系统识别出10种不同的转换模式（包括4种先前未记录的变种）。基于这些发现，我们提出了ChameleoScan，这是首个用于可靠验证变色龙应用的LLM驱动自动化UI探索框架。该系统通过其核心创新——预测性元数据分析、语义界面理解和类人交互策略，在保持本地决策可解释性的同时确保全局检测一致性。对1,644个iOS应用的全面评估展示了其操作效能（9.85%检测率，92.59%精确度），且研究结果已获Apple正式认可。实现代码和数据集可在</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/ChameleoScan" target="_blank">https://github.com/ChameleoScan</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">获取。</span></span></p><p cid="n212" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1906-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1906-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n214" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">52、Character-Level Perturbations Disrupt LLM Watermarks</span></span></p><p cid="n215" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLM）水印已成为一种有前景的版权保护、滥用预防和机器生成内容检测技术。它在LLM生成过程中注入可检测信号，使相应的检测器能够进行后续识别。为了评估水印方案的鲁棒性，现有研究通常采用水印移除攻击，旨在通过修改水印文本来擦除嵌入的信号。然而，我们揭示现有水印移除攻击并非最优，这导致了一种误解，即有效的水印移除要么需要大的扰动预算，要么需要攻击者具备强大的能力，例如对目标LLM或其水印检测器进行无限查询。对移除攻击能力的系统性审视以及更复杂技术的发展在很大程度上仍未得到充分探索。因此，现有水印方案的鲁棒性可能被高估。为了填补这一空白，我们首先形式化了LLM水印的系统模型，并描述了两种受限于对水印检测器访问有限的真实威胁模型。然后我们分析了不同类型的扰动在其攻击范围上的差异，即单次编辑能够影响的标记数量。我们观察到，字符级扰动（如拼写错误、交换、删除、同形异义字）通过破坏标记化过程可以同时影响多个标记。我们证明，在最严格的威胁模型下，字符级扰动相比标记级或句子级方法在移除水印方面显著更有效。我们进一步提出了基于遗传算法（GA）的引导式移除攻击，该算法使用参考检测器进行优化。在具有对水印检测器有限黑盒查询的实际威胁模型下，我们的方法展示了强大的移除性能。在五个代表性水印方案和两个广泛使用的LLM上的实验一致证实了字符级扰动的优越性以及参考检测器引导的GA在现实约束下移除水印的有效性。此外，我们认为在考虑潜在防御时存在一种对抗困境：任何固定防御都可以通过适当的扰动策略绕过。基于这一原则，我们提出了一种自适应复合字符级攻击。实验结果表明，这种方法可以有效防御现有防御。我们的研究突显了现有LLM水印方案中的重大漏洞，并强调了开发新型鲁棒机制的紧迫性。</span></span></p><p cid="n216" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s138-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s138-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n218" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">53、Characterizing the Implementation of Censorship Policies in Chinese LLM Services</span></span></p><p cid="n219" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">从中国防火墙实施的网络级审查，到TOM-Skype和微信等第三方服务实施的特定平台机制，中国的互联网审查一直在随着新技术的发展而不断演变。在当前的AI时代，像大语言模型（LLMs）这样的新兴工具也不例外。然而，确保符合中国严格的法定审查标准，对服务提供商来说是一项独特而复杂的挑战。虽然目前关于大语言模型内容审核的研究主要集中在对齐技术上，但这些技术缺乏可靠性，无法充分满足严格执行的信息管控要求。</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在这项工作中，我们首次对嵌入中文大语言模型（LLM）服务中的显性屏蔽进行了研究。我们利用活跃聊天会话期间服务器与客户端之间通信中的信息泄露，旨在找出屏蔽决策在LLM服务工作流程中的嵌入位置。我们观察到，百度文心一言、DeepSeek、豆包、Kimi和通义千问等知名服务持续依赖传统、过时的屏蔽策略。我们发现屏蔽设置在输入、输出和搜索阶段，后两个阶段会向客户端机器泄露不同数量的被审查信息，包括近乎完整的回复和未在浏览器中呈现的搜索参考内容。</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">鉴于有必要在全球舞台上的竞争与本土审查限制之间取得平衡，我们实时观察到托管模型的服务提供商在自我矛盾中做出的让步。通过这项工作，我们强调了构建更全面的大语言模型（LLM）内容可访问性威胁模型的重要性，该模型应整合实时部署，以研究与现实世界使用相关的访问情况，特别是在审查严格的地区。</span></span></p><p cid="n220" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1761-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1761-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n222" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">54、Chasing Shadows: Pitfalls in LLM Security Research</span></span></p><p cid="n224" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型(LLMs)在安全研究中日益普及。然而，它们的独特特性引入了一些挑战，这些挑战削弱了可重复性、严谨性和评估的既定范式。先前的工作已经确定了传统机器学习研究中的常见陷阱，但这些研究早于LLMs的出现。在本文中，我们确定了九种常见的陷阱，这些陷阱随着LLMs的出现而变得(更加)相关，并且可能损害涉及它们的研究的有效性。这些陷阱贯穿整个计算过程，从数据收集、预训练和微调到提示和评估。我们评估了这些陷阱在2023年至2024年间所有72篇发表在顶级安全和软件工程会议上的同行评审论文中的普遍性。我们发现每篇论文至少包含一个陷阱，且每个陷阱出现在多篇论文中。然而，只有15.7%的当前陷阱被明确讨论，表明大多数陷阱仍未被认识到。为了了解它们的实际影响，我们进行了四个实证案例研究，展示了个别陷阱如何误导评估、夸大性能或损害可重复性。基于我们的发现，我们提供了可行的指导方针，以支持社区未来的工作。</span></span></p><p cid="n225" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1749-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1749-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n227" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">55、Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation</span></span></p><p cid="n228" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">内部威胁可导致不可接受的损失，是一种普遍且重要的安全问题，因此其检测至关重要。近年来，基于机器学习的内部威胁检测（ITD）方法已被提出并取得了有前景的结果。尽管取得了这些成功，但一个主要挑战——数据不足——限制了这些ITD方法的进一步发展。矛盾之处在于，企业内部数据高度敏感且通常无法获取，而公共数据集要么在现实世界覆盖方面有限，要么在合成数据的情况下缺乏丰富的语义信息和真实的行为模式。因此，构建真实的内部威胁数据集至关重要。为应对这一挑战，我们提出了Chimera，这是首个基于大型语言模型（LLM）的多智能体框架，可自动模拟良性和恶意内部活动，并收集跨不同企业环境的日志。基于对组织构成和结构特征的分析，Chimera通过详细的角色建模定制每个LLM智能体以代表单个员工，并与小组会议、成对互动和自组织调度等模块相结合。通过这种方式，Chimera能够准确反映真实企业运营的复杂性。Chimera的当前版本包含15种不同类型的手工抽象内部攻击，如知识产权盗窃和系统破坏。使用Chimera，我们在三种典型的数据敏感型组织场景（包括科技公司、金融机构和医疗机构）中模拟良性和攻击活动，并生成了一个名为ChimeraLog的新数据集，以促进基于机器学习的ITD方法的发展。为评估ChimeraLog的质量和真实性，我们进行了全面的人类研究和定量分析。结果表明该数据集具有多样性和真实性。进一步的专业分析突显了真实威胁模式的存在以及可解释的活动轨迹。此外，我们在ChimeraLog上评估了现有内部威胁检测方法的有效性。平均F1得分为0.83，显著低于在基准数据集CERT上观察到的0.99分，从而说明了ChimeraLog在威胁检测任务中带来的更大难度。</span></span></p><p cid="n229" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f375-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f375-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n231" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">56、Cirrus: Performant and Accountable Distributed SNARK</span></span></p><p cid="n232" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">简洁非交互式知识论证（SNARKs）能够在许多应用中实现计算的高效验证。然而，为大规模任务（如可验证机器学习或虚拟机）生成SNARK证明在计算上仍然昂贵。一种有前景的方法是将证明生成工作负载分布在多个工作节点上。一个实用的分布式SNARK协议应具有三个特性：水平扩展性且开销低（每个工作节点线性计算和对数级通信）、可追责性（高效检测恶意工作节点）以及与电路和工作节点数量无关的通用可信设置。现有协议无法同时实现所有这些特性。在本文中，我们提出了Cirrus，这是首个同时实现所有三种理想特性的分布式SNARK生成协议。我们的协议基于HyperPlonk（EUROCRYPT&#39;23），继承了其通用可信设置。它实现了工作节点和协调器的线性计算复杂度，同时具有低通信开销。为实现可追责性，我们引入了一种高效的追责协议来定位恶意工作节点。此外，我们提出了一种分层聚合技术，以进一步减少协调器的工作负载。我们在硬件适中的机器上实现并评估了Cirrus。实验表明，Cirrus具有高度可扩展性：使用32台8核机器，在40秒内即可为拥有3300万门电路的证明生成。与最先进的可追责协议Hekaton（CCS&#39;24）相比，Cirrus在PLONK友好型电路（如Pedersen哈希）上的证明生成速度提高了7倍以上。我们的追责协议也能在4秒内高效识别出故障工作节点，使Cirrus特别适用于去中心化和外包计算场景。</span></span></p><p cid="n233" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f668-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f668-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n235" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">57、CoLD: Collaborative Label Denoising Framework for Network Intrusion Detection</span></span></p><p cid="n236" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">标签噪声在网络入侵检测中构成了重大挑战，导致错误分类和检测准确率下降。处理噪声标签的现有方法通常缺乏对网络流量的深入洞察，盲目重建标签分布以过滤带有噪声标签的样本，从而造成次优性能。本文从因果关联的角度揭示了噪声标签对入侵检测模型的影响，将性能下降归因于网络流量中跨类别的局部特征一致性。受此启发，我们提出了CoLD，一个用于网络入侵检测的协同标签去噪框架。CoLD将原始特征集划分为多个子集，采用局部联合学习来破坏局部一致性，迫使编码器学习细粒度和鲁棒的表示。它进一步应用因果协同去噪，通过分析多种表示与其潜在真实标签之间的因果差异来检测和过滤噪声标签，从而生成一个经过净化的数据集用于训练抗噪声分类器。在多个基准数据集上的实验表明，CoLD有效提升了分类性能和对标签噪声的鲁棒性，凸显了其在增强嘈杂环境中网络入侵检测系统的潜力。</span></span></p><p cid="n237" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1950-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1950-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n239" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">58、Connecting the Dots: An Investigative Study on Linking Private User Data Across Messaging Apps</span></span></p><p cid="n241" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">移动消息应用已成为日常交流的重要组成部分，拥有庞大的用户基础（例如，Telegram超过9.5亿用户，KakaoTalk达4870万用户）。为提升用户参与度和扩大用户规模，消息应用提供了丰富多样的上下文相关和平台特定功能，如附近用户搜索、联系人发现以及基于单点登录（SSO）的账户链接。虽然这些功能使用户能够在单个移动设备上使用多种消息应用，但它们也带来了跨多个消息应用链接私人用户信息的隐私风险，这一问题尚未得到充分研究。本文对韩国广泛使用的消息应用（包括KakaoTalk、Telegram、WhatsApp、Signal和Tinder）中的隐私威胁进行了深入分析，展示了利用联系人发现、基于SSO的账户链接和附近用户搜索功能的具体攻击实例，这些攻击会损害用户隐私。更重要的是，我们将这些攻击串联起来，实施了首个跨平台链接攻击，使攻击者能够去匿名化用户名，并推断大量非目标用户和目标用户的物理位置，平均误差范围为324米。我们的研究结果表明，保障联系人发现的安全性至关重要，因为宽松的联系人发现政策允许攻击者利用电话号码和个人资料图片作为链接键，跨多个消息应用连接私人用户信息。我们讨论并提出了缓解策略以减轻所呈现的威胁。</span></span></p><p cid="n242" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s556-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s556-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n244" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">59、Consensus in the Known Participation Model with Byzantine Faults and Sleepy Replicas</span></span></p><p cid="n245" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们研究了已知参与模型中的一致性问题，该模型同时存在拜占庭故障和休眠副本，其中诚实副本可能不可预测地进入休眠状态，且副本知道最少活跃诚实副本的数量。我们的主要贡献是对这种混合故障模型中的一致性问题进行了细粒度处理。首先，我们提出了一个同步原子广播协议，其期望延迟为$5Delta+2delta$，最佳情况延迟为$2Delta+2delta$，其中$Delta$是网络延迟的上界，$delta$是实际网络延迟。其次，在部分同步网络中（$Delta$值未知），我们表明可以使传统的拜占庭容错(BFT)协议容忍休眠副本，但必须做出稳定存储假设（副本需要将中间共识参数存储在稳定存储中）。最后，在部分同步网络但不假设稳定存储的情况下，我们展示了关于副本总数$n$、拜占庭副本最大数量$f$和同时休眠副本最大数量$s$之间关系的几个界限。利用这些界限，我们将HotStuff (PODC&#39;19)转化为一个能够容忍休眠副本而不牺牲性能的协议。</span></span></p><p cid="n246" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s448-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s448-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n248" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">60、Constructive Noise Defeats Adversarial Noise: Adversarial Example Detection for Commercial DNN Services</span></span></p><p cid="n249" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">商业深度神经网络服务已以机器学习即服务（MLaaS）的形式发展起来。为缓解对抗样本的潜在威胁，已提出了各种检测方法。然而，现有方法通常需要访问目标模型的细节或训练数据集，这在MLaaS场景中通常不可用。在无法获取目标模型细节或训练数据集的情况下，这些方法的检测准确率会显著下降。在本文中，我们提出了Falcon，一种由第三方提供的对抗样本检测方法，能够同时实现准确性和效率。基于干净样本和对抗样本在噪声容忍度上的差异，我们探索了一种建设性噪声，这种噪声添加到干净样本中不会影响模型的输出标签，但当添加到对抗样本中时，会导致模型输出发生明显变化。对于每个输入，Falcon生成具有特定分布和强度的建设性噪声，并通过添加建设性噪声前后目标模型输出的差异来实现检测。我们在4个公共数据集上进行了大量实验，以评估Falcon在检测10种典型攻击时的性能。Falcon优于最先进的检测方法，实现了对抗样本的最高真阳性率（TPR）和干净样本的最低假阳性率（FPR）。此外，Falcon在6个知名商业深度神经网络服务上实现了约80%的TPR和5%的FPR，性能优于最先进的方法。即使对手完全了解检测细节，Falcon也能保持其准确性。</span></span></p><p cid="n250" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s250-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s250-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n252" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">61、Continuous User Behavior Monitoring using DNS Cache Timing Attacks</span></span></p><p cid="n253" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">域名系统（DNS）是互联网的核心组成部分。客户端查询DNS服务器以将域名转换为IP地址。本地DNS缓存减少了查询DNS服务器所需的时间，从而降低了连接尝试的延迟。先前的研究表明，可以通过时序攻击利用DNS缓存来测试用户最近是否访问过特定网站，但这些研究缺乏驱逐功能，即无法精确监控用户访问网站的时间。其他研究则专注于路由器中的DNS缓存。所有先前的攻击都需要在受害者系统上执行某种形式的代码（例如原生代码、Java或JavaScript），而这并不总是可行的。我们引入了DMT，这是一种新颖的Evict+Reload攻击，可通过本地系统范围的DNS缓存持续监控受害者的互联网访问。DMT的基础是可靠的DNS缓存驱逐：我们提出了4种DNS缓存驱逐技术，用于在无权限和沙盒化原生攻击、虚拟化跨VM攻击以及基于浏览器的攻击（即带有JavaScript的网站和利用网站中字体串行加载的无脚本攻击）中驱逐本地DNS缓存。我们的攻击在默认设置以及使用DNS-over-TLS、DNSSEC或非默认DNS转发器进行安全防护时均有效。在我们的最快驱逐原语下，我们在所有上下文中观察到的平均驱逐时间为77.267毫秒，重新加载和测量时间在最佳情况（跨VM攻击）下为100个域名平均685.86毫秒，在最坏情况（基于JavaScript的攻击）下平均14.710秒。因此，对于五分钟粒度的攻击盲区，在最佳情况下小于0.26%，在最坏情况下为4.92%，这构成了可靠的攻击。在端到端的跨VM攻击中，我们可以在不到一秒的时间内可靠地检测出从103个网站列表（在开放世界场景中）的访问，F1得分为92.48%。在我们的基于JavaScript的攻击中，对于检测10个网站的访问，在有和无DNSSEC的情况下，我们分别实现了82.86%和78.89%的F1分数。我们认为DMT泄露了对敲诈和诈骗活动有价值的信息，或可用于提供针对受害者EDR解决方案的定制化漏洞利用。</span></span></p><p cid="n254" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2287-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2287-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n256" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">62、Convergent Privacy Framework for Multi-layer GNNs through Contractive Message Passing</span></span></p><p cid="n257" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">差分隐私(DP)已被集成到图神经网络(GNN)中，以保护敏感的结构信息，例如各种应用中的边、节点及相关特征。一种突出的方法是扰动消息传递过程，这是大多数GNN架构的核心。然而，现有方法通常会导致隐私成本随层数线性增长(例如，在Usenix Security&#39;23上发表的GAP)，最终需要添加过多噪声以维持合理的隐私水平。当使用表现优于单层GNN的多层GNN处理包含敏感信息的图数据时，这一局限性尤为突出。在本文中，我们通过将隐私放大技术应用于消息传递过程，并利用标准GNN操作固有的收缩特性，从理论上证明了隐私预算随层数收敛。受此分析启发，我们提出了一种简单而有效的收缩图层(CGL)，它在确保理论保证所需收缩性的同时保留了模型效用。我们的框架CARIBOU支持训练和推理，配备了收缩聚合模块、隐私分配模块和隐私审计模块。实验评估表明，CARIBOU显著改善了隐私-效用权衡，并在隐私审计任务中取得了卓越的性能。</span></span></p><p cid="n258" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f255-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f255-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n260" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">63、CoordMail: Exploiting SMTP Timeout and Command Interaction to Coordinate Email Middleware for Convergence Amplification Attack</span></span></p><p cid="n261" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">本文介绍了一种名为COORDMAIL的新型且强大的邮件汇聚放大攻击。传统的邮件DoS攻击主要向目标邮箱发送垃圾邮件，对邮件服务器运行的影响有限。相比之下，COORDMAIL利用SMTP协议的固有特性，即长会话超时和客户端控制的交互，巧妙地协调来自各种邮件中间件的反射邮件，最终将它们同时定向到入站邮件服务器。因此，不同邮件中间件的放大能力被集中起来，形成高度放大的攻击流量。从SMTP会话状态机和邮件反射行为出发，我们确定了众多适用于COORDMAIL的现实世界邮件中间件，包括10,079个反弹服务器、584个开放邮件中继和6个邮件转发服务提供商。通过构建SMTP命令序列，COORDMAIL能够以极低的速率与这些中间件保持长时间的SMTP通信，并控制它们在任何给定时刻稳定地反射邮件。我们证明COORDMAIL以低成本高效：1000个SMTP连接可实现超过30,000倍的带宽放大。虽然大多数现有安全机制对COORDMAIL无效，但我们提出了可行的缓解措施，可将COORDMAIL的汇聚放大能力降低数十倍。我们已负责任地向邮件中间件和主流邮件服务提供商报告了COORDMAIL，其中一些已接受我们的建议。</span></span></p><p cid="n262" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1414-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1414-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n264" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">64、CoT-DPG: A Co-Training based Dynamic Password Guessing Method</span></span></p><p cid="n265" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">密码仍然是主要的身份验证方法，安全界研究密码猜测以增强密码安全性。动态密码猜测持续收集目标信息并在猜测过程中动态拟合分布，从而扩大了威胁。现有方法主要分为两类：动态调整密码策略和基于生成模型的动态生成。然而，这些方法从单一视角拟合目标分布，忽略了不同维度信息之间的互补效应。如果能充分利用多维度信息，动态密码猜解性能将大幅提升，但如何有效融合多维度信息仍是一个挑战。受此启发，我们提出了CoT-DPG，一种新型动态密码猜解框架，允许多个猜解模型协作学习并互补知识。这是协同训练方法在多视图学习中首次应用于密码猜解。首先，在特征层面，我们基于增量训练动态更新神经网络参数并拟合目标分布。其次，在字符层面，我们设计了策略分布优化方法以减轻策略选择的盲目性。第三，我们采用协同训练方法进行多维度互补学习、迭代训练和密码生成。最后，实验证明了所提框架的有效性，在八个真实世界密码数据集上，与最先进方法相比，破解率绝对提升了6.4%至26.7%。</span></span></p><p cid="n266" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s755-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s755-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n268" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">65、Crack in the Armor: Underlying Infrastructure Threats to RPKI Publication Point Reachability</span></span></p><p cid="n269" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">RPKI在防止BGP前缀劫持方面的有效性不仅依赖于有效ROA的存在，还依赖于依赖方（RPs）从发布点（PPs）成功检索ROA的能力。在此检索过程中保证数据完整性和不间断连接，需要正确实施底层基础设施（即DNS和路由基础设施）中的安全措施。在本文中，我们收集了信息检索过程中使用的具体DNS和路由基础设施信息，并分析了影响RPKI PP可达性的基础设施威胁。关于DNS基础设施，我们报告显示31个PP（48.4%）容易受到DNS欺骗攻击，并指出了DNSSEC未保护区域出现的原因，例如重定向到未保护区域的CNAME和委托给第三方不安全DNS服务器的NS记录。关于与名称服务器通信的路由基础设施，我们的分析显示，多达55个PP（85.9%）在其解析路径上至少有一个未受ROA保护的名称服务器，并强调gTLD名称服务器缺乏ROA注册是其中44个PP存在漏洞的原因。关于RP-PP通信的路由基础设施，我们报告有5个PP未为其PP服务器的IP地址注册ROA。路由劫持攻击的模拟表明，在最脆弱的PP情况下，高达65%到83%的自治系统（ASes）可能会失去与该PP的连接。此外，我们研究了发布点之间的确定性和概率性依赖关系，发现了一个关键问题：一些由RIR运营的PP依赖于安全性较低的下层PP，这会显著放大不安全PP中的漏洞影响，可能导致级联故障。</span></span></p><p cid="n270" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1141-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1141-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n272" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">66、CRISP: An Efficient Cryptographic Framework for ML Inference Against Malicious Clients</span></span></p><p cid="n273" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于半诚实安全模型的机器学习推理协议在实际应用中容易受到恶意客户端的攻击，这些攻击可能导致机器学习模型参数的泄露。先前的研究引入了额外的MAC计算来确保客户端行为的正确性，但这在线上推理阶段增加了运行时间和通信成本。在本工作中，我们提出了CRISP，一个高效的两方密码学框架，旨在防御恶意客户端的攻击。具体而言：1）我们基于一种新的密码学原语（函数秘密共享）设计了非线性层的协议，我们方法的核心是优化MAC的重构过程。2）我们为线性层提出了一个复数域验证机制，该机制通过更好地利用同态加密CKKS中的复数空间，消除了额外的MAC计算。此外，在我们之前的工作（SIMC，USENIX Security&#39;22）中，我们识别了实际应用中的兼容性问题。当应用某些混淆电路优化时，非线性层中的MAC重构过程可能会泄露模型的中间输入和输出。相比之下，CRISP有效地避免了这一问题。在SIMC考虑的安全推理基准测试中，CRISP将机器学习推理的总通信成本降低了高达94%，并将推理延迟减少了高达43%。</span></span></p><p cid="n274" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s11-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s11-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n276" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">67、Cross-Boundary Mobile Tracking: Exploring Java-to-JavaScript Information Diffusion in WebViews</span></span></p><p cid="n277" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">WebViews是将基于Web的内容嵌入Android应用的常用方法。尽管它们提供与浏览器类似的功能并在隔离环境中执行，但应用可以通过在运行时动态注入JavaScript代码直接干扰WebViews。尽管先前工作已广泛分析了应用的Java代码，但现有框架对WebView中执行的JavaScript代码的可见性有限。因此，人们对WebView中执行的脚本的行为和特征以及是否存在隐私违规行为的理解有限。为解决这一差距，我们提出了WebViewTracer，这是一个旨在在运行时动态分析WebView中JavaScript代码执行情况的框架。我们的系统将WebView内部的JavaScript执行跟踪与Java方法调用信息相结合，以捕获Java SDK和Web脚本之间发生的信息交换。我们利用WebViewTracer对10K个Android应用数据集进行了首次大规模的WebView内部隐私违规行为动态分析。我们检测到4,597个加载WebView的应用，发现其中超过69%的应用将敏感和跟踪相关信息（通常是JavaScript代码无法访问的信息）注入到WebView中。这包括广告ID和Android构建ID等标识符。关键的是，90%的应用使用基于Web的API将这些信息泄露给第三方服务器。我们还发现了WebView中的JavaScript代码使用常见的Web指纹识别技术的具体证据，这些技术可以补充其跟踪信息。我们观察到，WebView的动态特性正在被积极利用，以便在移动跟踪生态系统中的多个参与者之间扩散敏感信息，这表明Android WebView存在隐私风险。通过揭示这些持续的隐私违规行为，我们的研究旨在促使平台利益相关者对嵌入式Web技术的使用进行更多审查，并强调需要额外的安全措施。</span></span></p><p cid="n278" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s910-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s910-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n280" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">68、Cross-Cache Attacks for the Linux Kernel via PCP Massaging</span></span></p><p cid="n282" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">尽管内存损坏防御研究已有数十年历史，内核内存分配器仍然是一个关键的攻击面。尽管最近的缓解策略降低了传统攻击技术的有效性，但我们证明稳健的跨缓存攻击仍然可行并构成重大威胁。在本文中，我们介绍了PCPLost，一种跨缓存内存按摩技术，它通过巧妙利用侧信道推断内核分配器的内部状态来绕过主流缓解措施。我们证明，诸如越界(OOB)漏洞——以及通过支点利用的释放后使用(UAF)和双重释放(DF)漏洞——可以通过跨缓存攻击可靠地利用，适用于所有通用缓存，即使在存在噪声的情况下也是如此。我们通过利用PCPLost利用6个公开披露的CVE漏洞，验证了我们方法的通用性和稳健性，并讨论了可能的缓解措施。我们的方法在获取跨缓存布局方面具有显著的可靠性（大多数情况下超过90%），这表明当前的缓解策略无法在Linux内核中为此类攻击提供全面保护。</span></span></p><p cid="n283" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f862-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f862-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n285" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">69、Cross-Consensus Reliable Broadcast and its Applications</span></span></p><p cid="n286" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">传统的拜占庭容错共识协议主要关注节点组内部的工作流程。近年来，许多共识应用涉及跨组通信。例如，不同基础设施上复制状态机之间的通信、基于分片的协议中不同分片节点之间的通信，以及跨链桥接。然而，很少有人致力于建模跨组通信的属性。在这项工作中，我们提出了一种名为跨共识可靠广播（XRBC）的新原语。XRBC原语建模了两个组之间通信的安全属性，其中至少有一个组执行共识协议。我们在不同假设下提供了三种XRBC构造，并展示了三种不同的XRBC协议应用：通过Reticulum（NDSS 2024）的案例研究实现的跨分片协调协议，通过Chainspace（NDSS 2018）的案例研究实现的跨分片交易协议，以及跨链桥接解决方案。我们的评估结果表明，我们的协议具有很高的效率，并能惠及不同的应用。例如，在我们对Reticulum的案例研究中，我们的方法比传统方法实现了61.16%的更低延迟。</span></span></p><p cid="n287" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s207-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s207-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n289" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">70、Cryptobazaar: Private Sealed-bid Auctions at Scale</span></span></p><p cid="n290" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">本文介绍了Cryptobazaar，一种可扩展、私密且去中心化的密封投标拍卖协议。特别地，我们的协议通过保护未中标出价者的出价机密性，同时确保结果的可公开验证性，并仅依赖单个不可信拍卖师进行协调，从而保护未中标出价者的隐私。Cryptobazaar的核心是将一个用于计算一元编码出价列表逻辑或的高效分布式协议，与多种新颖的零知识简洁知识论证相结合，这些论证可能具有独立的学术价值。我们提出了协议的多种变体，可用于高效进行第一价格、第二价格以及更一般的(p+1)价格拍卖，以及顺序第一价格拍卖。最后，我们对Cryptobazaar实现的性能评估表明该协议具有高度实用性。例如，一次包含128名出价者、价格范围为1024个值的拍卖在0.5秒内完成，且每个出价者仅需发送和接收约32KB的数据。</span></span></p><p cid="n291" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f481-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f481-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n293" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">71、CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning</span></span></p><p cid="n294" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">公开可用的预训练大模型（即主干网络）和用于参数高效微调（PEFT）的轻量级适配器已成为现代机器学习流水线的标准组件。然而，在推理过程中保护用户输入和微调适配器的隐私（这些适配器通常在敏感数据上训练）仍然是一个重大挑战。将密码学技术（如多方计算（MPC））应用于PEFT设置仍然会在主干网络和适配器之间产生大量加密计算，这主要是由于它们之间固有的双向通信。为解决这一限制，我们提出了CryptPEFT，这是首个专为私有推理场景设计的PEFT解决方案。CryptPEFT引入了一种新颖的单向通信（OWC）架构，将加密计算仅限制在适配器内，显著降低了计算和通信开销。为在此约束下保持强大的模型效用，我们探索了OWC兼容适配器的设计空间，并采用自动化架构搜索算法来优化私有推理效率与模型效用之间的权衡。我们在广泛使用的图像分类数据集上使用Vision Transformer主干网络对CryptPEFT进行了评估。结果表明，CryptPEFT显著优于现有基线，在模拟广域网（WAN）和局域网（LAN）环境中实现了20.62倍至291.48倍的加速。在CIFAR-100上，CryptPEFT仅需2.26秒的推理延迟即可达到85.47%的准确率。这些研究结果表明，CryptPEFT为现代基于PEFT的推理提供了一种高效且隐私保护的解决方案。</span></span></p><p cid="n295" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1102-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1102-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n297" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">72、CTng: Secure Certificate and Revocation Transparency</span></span></p><p cid="n298" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们提出了CTng，这是一种演进而实用的PKI设计，能有效解决现有PKI系统面临的多项关键挑战。CTng确保了强大的安全特性，包括证书的透明保证和明确无误的撤销保证，这些都是在NTTP安全模型下实现的，即无需信任任何单一的CA、日志记录方或依赖方。即使在这些实体存在任意腐败的情况下，这些保证仍然成立，只需假设腐败监控者的数量有一个已知上限（例如f=8），且对性能的影响最小。CTng还支持离线证书验证并保护依赖方的隐私，同时提供可扩展且高效的撤销更新分发。这些特性显著优于当前的PKI设计。特别是，虽然证书透明（CT）旨在消除单一信任点，但现有规范仍然假设日志记录方是善意的。通过日志冗余来解决这个问题是可能的，但效率较低，限制了部署配置中f≤2的情况。我们提供了对CTng开源原型（安全分析和评估），表明它在实际部署条件下是高效且可扩展的。</span></span></p><p cid="n299" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s213-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s213-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n301" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">73、CtPhishCapture: Uncovering Credential-Theft-Based Phishing Scams Targeting Cryptocurrency Wallets</span></span></p><p cid="n302" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">由于涉及巨大的经济利益，基于凭证窃取的加密货币钱包网络钓鱼（CtPhish）骗局已成为加密货币生态系统中最为普遍的恶意活动之一。在这些攻击中，受害者被诱骗访问CtPhish网站或应用程序，并受骗泄露其凭证，从而使攻击者能够窃取其加密货币资产。尽管存在几种网络钓鱼检测方法，但它们要么不适用于CtPhish，要么存在显著局限性。为填补这一空白，我们提出了CtPhishCapture，一个针对CtPhish网站和应用程序的大规模检测系统。CtPhishCapture访问可疑网站，采用基于大型语言模型（LLM）的检测方法来识别CtPhish网站，并尝试下载和分析潜在的CtPhish应用程序以进行进一步检测。经过六个月的部署，CtPhishCapture识别出5,138个CtPhish网站和10,612个CtPhish应用程序。值得注意的是，只有17%的网站和21%的应用程序先前被社区报告过，这表明CtPhishCapture新发现了83%的网站和79%的应用程序，使其成为迄今为止已知最大的CtPhish检测系统。利用收集的数据集，我们对CtPhish生态系统进行了全面的端到端测量和分析。我们的分析研究了攻击者如何诱骗受害者访问CtPhish网站和应用程序，如何获取用户信任，以及最终如何窃取受害者的加密货币资产。此外，我们还对相关网站和应用程序进行了深入测量，包括其特征、规避技术和估计的财务损失。最后，我们与一家领先的搜索引擎提供商合作部署了CtPhishCapture。通过整合CtPhishCapture的检测结果，每周关于CtPhish的用户投诉减少了5.8倍。</span></span></p><p cid="n303" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2854-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2854-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n305" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">74、cwPSU: Efficient Unbalanced Private Set Union via Constant-weight Codes</span></span></p><p cid="n306" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">私有集合并（PSU）允许两方在不泄露任何额外信息的情况下计算其私有集合的并集。尽管已有几种针对非平衡场景的PSU协议被提出，但随着较大集合规模的增加，这些构造方案仍存在显著的通信开销。此外，它们对无意识伪随机函数的多重调用导致通信轮次增加，这已成为实际应用中的瓶颈。在本工作中，我们提出了cwPSU，一种基于常重码和层次全同态加密的新型非平衡PSU协议。为防止信息泄露，我们引入了一种称为批量密文重排的新技术，实现了打包密文的安全重排序。此外，我们提出了一种优化的算术常重等价算子，将非标量乘法的数量减少到朴素方法所需的三分之一。我们协议的通信复杂度与较小集合的大小呈线性关系，且与较大集合的大小无关。值得注意的是，cwPSU仅需一轮在线通信。实验结果表明，cwPSU在各种网络条件下均优于现有最先进协议，实现了通信量减少5.1至32.4倍，运行时间加速3.1至13.3倍。</span></span></p><p cid="n307" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1128-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1128-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n309" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">75、Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction Attack</span></span></p><p cid="n310" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">为保护高价值深度神经网络的知识产权，黑盒水印技术已成为一种关键防御手段并日益受到重视。这些方法通过精心设计的触发样本将水印嵌入到模型的预测行为中，从而能够通过API查询进行验证。同时，模型提取攻击通过利用查询访问来复制带水印的模型，从而威胁专有深度学习模型。这些攻击也为水印方案的稳健性和对抗能力提供了见解。然而，先前的方法难以去除水印信息，无意中保留了防御机制。它们还存在效率低下的问题，通常需要数千次查询才能达到竞争性能。为解决这些局限性，我们提出了一个名为SSLExtraction的查询高效模型提取框架。SSLExtraction通过特征空间中的贪婪随机游走选择查询，从而实现有效的模型复制和水印去除。具体而言，SSLExtraction遵循自监督学习范式提取内在数据表示，将原始像素级输入转换为与水印无关的特征。然后，我们在特征空间中提出了一种贪婪随机游走算法，以构建一个分布良好的查询集，有效覆盖特征空间同时避免冗余查询。通过在特征空间中选择查询，我们的方法自然地将水印模式识别为异常值，从而实现同时去除水印。此外，我们提出了一种专门为水印任务设计的评估指标，强调良性模型与被盗模型之间的区别。与依赖手动预定义阈值的前期方法不同，我们的评估指标采用假设检验来衡量可疑模型与带水印模型和良性模型之间的相对距离，识别可疑模型最接近的模型。实验结果表明，与基线方法相比，我们的方法显著降低了查询成本，同时在各种数据集和水印场景中有效去除了水印。</span></span></p><p cid="n311" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f223-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f223-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n313" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">76、Decompiling the Synergy: An Empirical Study of Human–LLM Teaming in Software Reverse Engineering</span></span></p><p cid="n315" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）正在改变以往由人类主导的领域。本研究首次系统性地探讨了LLMs如何在软件逆向工程（SRE）过程中与分析师协作。为此，我们首先通过一项针对153名从业者的在线调查，记录了LLMs在SRE领域的应用现状，然后设计了一项细粒度的人类研究，研究对象是两个具有代表性的真实世界软件的&#34;夺旗&#34;风格二进制文件。在我们的研究中，我们对48名参与者（分为24名新手和24名专家）的SRE工作流程进行了监测，观察了超过109小时的SRE过程。通过18项研究发现，我们揭示了LLMs在SRE中的各种益处和危害。值得注意的是，我们发现LLM辅助缩小了专业知识差距：新手的理解率提高了约98%，达到专家水平，而专家则获益甚微；然而，LLMs也会产生有害的幻觉、无用的建议和无效的结果。已知算法函数的筛选速度提高了2.4倍，工件恢复（名称、注释、类型）增加了至少66%。总体而言，我们的研究结果确定了人类与LLMs在SRE中的强大协同效应，但也强调了当前LLMs集成中的显著缺陷。</span></span></p><p cid="n316" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f380-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f380-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n318" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">77、Demystifying RPKI-Invalid Prefixes: Hidden Causes and Security Risks</span></span></p><p cid="n319" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">资源公钥基础设施（RPKI）通过利用路由源授权（ROA）对象将IP前缀与其合法的源ASN关联起来，从而增强互联网路由安全性。尽管RPKI部署迅速——目前已有超过51.3%的互联网路由被ROA覆盖，但截至今天仍有6,802个RPKI无效前缀。本研究首次对RPKI无效前缀的隐藏原因进行全面研究和分类，揭示ROA配置错误通常发生在IP租赁和IP传输服务过程中。我们确定了导致这些配置错误的场景，并将96.9%的RPKI无效前缀归因于此类配置错误。我们进一步展示了它们对数据平面的级联影响，指出虽然大多数前缀的影响可以忽略不计，但3.1%的前缀会导致完全连接丢失，7.1%的前缀通过增加延迟和额外跳数来降低路由性能——在某些情况下甚至会绕过预期的安全机制；此外，我们发现此类配置错误正在触发劫持检测系统的误报。为验证我们的研究结果，我们通过与174个网络运营商直接合作，构建了一个包含294个配置错误前缀的真实数据集。我们还采访了16家大型ISP和主要租赁经纪人关于其ROA管理实践，并提出了避免ROA配置错误的建议。总之，这项研究不仅填补了先前研究的空白，还为网络运营商提供了可操作的改进ROA管理和减少RPKI无效公告发生的建议。</span></span></p><p cid="n320" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s161-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s161-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n322" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">78、Demystifying the Access Control Mechanism of ESXi VMKernel</span></span></p><p cid="n324" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">VMware ESXi是一种广泛部署的企业级1型管理程序，作为现代云基础设施的基础。为加强特权隔离，ESXi在VMKernel中引入了强制访问控制机制。然而，由于VMKernel的专有和闭源特性，其内部访问控制架构在很大程度上仍不透明且未被充分探索。先前的研究主要集中在虚拟设备漏洞和虚拟机逃逸上，而VMKernel的内部访问控制机制和特权模型则很少被检查。为填补这一空白，我们对VMKernel的访问控制机制进行了首次全面的安全分析。我们开发了一种面向域-控制结构的分析方法来重建关键内部权限逻辑，并设计了一种结构感知的调试框架以支持细粒度的运行时验证。利用该框架，我们发现了几个关键的设计缺陷，包括可写且不受保护的内存控制结构以及可被利用的开发者保留的系统调用接口。我们演示了三种实际攻击场景，这些场景利用这些缺陷来绕过沙箱限制、提升权限并获得持久访问。总之，我们向VMware报告了14个漏洞，所有漏洞均已得到确认和修复，共获得42,000美元的漏洞赏金。</span></span></p><p cid="n325" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f700-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f700-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n327" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">79、DirtyFree: Simplified Data-Oriented Programming in the Linux Kernel</span></span></p><p cid="n328" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着内核控制流完整性（KCFI）的出现，面向数据编程（DOP）已成为传统控制流劫持技术（如返回导向编程，ROP）的重要替代方案。与控制流攻击不同，DOP通过操作内核数据流实现权限提升，而无需违反控制流完整性。然而，传统的DOP攻击由于其多阶段特性仍然复杂且实用性有限，通常需要堆地址泄露、任意地址读取和任意地址写入能力。每个阶段都对内核对象的选择和使用施加了严格限制。为解决这些限制，我们引入了DirtyFree，这是一种利用任意释放原语的系统性利用方法。该原语能够强制释放攻击者控制的内核对象，显著降低利用要求并简化整体利用过程。DirtyFree提供了一种在多种内核缓存中识别合适的任意释放对象的系统方法，并提出了针对安全关键对象（如cred）的结构化利用策略。通过广泛评估，我们成功识别出覆盖大多数内核缓存的14个任意释放对象，通过成功利用24个真实世界内核漏洞证明了DirtyFree的实际有效性。此外，我们提出并实现了两种旨在缓解DirtyFree的缓解技术，有效防止了利用，同时仅产生微不足道的性能开销（分别为0.28%和-0.55%）。</span></span></p><p cid="n329" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f527-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f527-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n331" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">80、Discovering Blind-Trust Vulnerabilities in PLC Binaries via State Machine Recovery</span></span></p><p cid="n332" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可编程逻辑控制器（PLC）是控制具有现实世界物理效应设备的工业计算机，这些系统中的安全漏洞可能导致灾难性后果。尽管先前的研究已提出检测PLC状态机中安全问题的技术，但大多数方法需要访问设计规范或源代码——这些资源通常对分析师或终端用户不可用。本文针对一类普遍存在的漏洞，我们将其命名为&#34;盲目信任漏洞&#34;，这些漏洞由外围输入上缺失或不完整的安全检查引起。我们引入了Ta&#39;veren，这是一个新颖的基于静态分析的框架，可以直接从PLC二进制文件中识别此类漏洞，而不依赖于固件重托管，这仍然是固件分析中的一个开放研究问题。Ta&#39;veren恢复了PLC二进制文件中的有限状态机，从而能够在各种规范下重复进行安全分析。为了将程序状态抽象为逻辑相关状态，我们利用了PLC一致使用特定变量表示内部状态的见解，从而允许进行激进的状态去重。这一见解使我们能够在不损害完备性的情况下有效去重状态。我们开发了Ta&#39;veren的原型并在真实的PLC二进制文件上对其进行了评估。实验表明，Ta&#39;veren能够高效地恢复有意义的有限状态机，并以高有效性发现关键的安全违规。</span></span></p><p cid="n333" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1624-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1624-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n335" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">81、Distributed Broadcast Encryption for Confidential Interoperability across Private Blockchains</span></span></p><p cid="n336" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">跨分布式账本技术(DLT)网络的互操作依赖于账本状态从一个网络到另一个网络的安全传输。对于访问权限仅限于注册成员的私有网络而言，这尤其具有挑战性。现有方法依赖于一个可信的集中式代理，该代理接收一个网络的加密账本状态，解密它，然后将其发送到另一个网络的成员。尽管这种方法有效，但它违背了DLT的基本原则，即避免单点故障（或单一信任源）。在本文中，我们利用全分布式广播加密(FDBE)构建了一个用于私有网络间机密信息共享的完全去中心化协议。与传统广播加密(BE)相比，FDBE的特点是分布式设置和密钥生成，即互不信任的各方无需可信设置即可就BE的公钥达成一致，并安全地派生其解密密钥。给定任何FDBE，两个私有网络可以安全地共享信息：一个网络中的发送者使用另一个网络的FDBE公钥为其成员加密消息。所构建的方案在简化的通用可组合性(UC)框架下是安全的。为进一步证明我们方法的实用性，我们提出了首个具有恒定大小解密密钥和密文的FDBE实例，并通过一个参考实现评估了其性能，该实现考虑了Hyperledger Cacti互操作框架内的两个私有Hyperledger Fabric网络。</span></span></p><p cid="n337" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1200-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1200-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n339" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">82、DNN Latency Sequencing: Extracting DNN Architectures from Intel SGX Enclaves with Single-Stepping Attacks</span></span></p><p cid="n341" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">深度神经网络（DNN）是现代计算的核心组成部分，支撑着图像识别、自然语言处理和音频分析等应用。这些模型的架构（例如，图层的数量和类型）被视为宝贵的知识产权，因为其设计需要大量的专业知识和计算投入。尽管可信执行环境（TEEs）如Intel SGX已被采用来保护这些模型，但最近关于模型提取攻击的研究表明，侧信道攻击（SCAs）仍可被用来提取DNN模型的架构。然而，许多现有的模型提取攻击要么没有考虑TEE的保护，要么仅限于特定类型的模型，降低了它们的实际应用性。在本文中，我们介绍了DNN延迟排序（DLS），这是一种新颖的模型提取攻击框架，针对在Intel SGX enclave中运行的DNN架构。DLS采用SGX-Step对模型执行单步操作并收集细粒度延迟轨迹，然后在函数和基本块级别进行分析以重建模型架构。我们的关键见解是，DNN架构本质上会影响执行行为，从而能够从延迟模式中实现准确的重建。我们在使用三种广泛使用的深度学习库（Darknet、TensorFlow Lite和ONNX Runtime）构建的模型上评估了DLS，并分别实现了97.3%、96.4%和93.6%的架构恢复准确率。我们进一步证明了DLS能够实现高级攻击，突显了其实用性和有效性。</span></span></p><p cid="n342" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1455-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1455-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n344" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">83、DOM-XSS Detection via Webpage Interaction Fuzzing and URL Component Synthesis</span></span></p><p cid="n345" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于DOM的跨站脚本（DOM-XSS）是一种普遍存在的Web漏洞。先前关于此类漏洞的大规模自动化检测和确认工作存在若干局限性。首先，先前的研究不与页面交互，因此无法执行依赖于用户操作的事件处理程序中的漏洞。其次，先前的研究无法找到URL组件，如GET参数和片段值，这些组件在用特定键/值实例化时会执行更多代码路径。为此，我们引入了SWIPE，这是一种DOM-XSS分析基础设施，它使用模糊测试生成用户交互以触发事件处理程序，并利用动态符号执行（DSE）自动合成URL参数和片段。我们在来自Tranco前30,000个热门域名的页面中找到的44,480个URL上运行了SWIPE。与先前的工作相比，SWIPE的模糊测试工具发现了多15%的漏洞。此外，我们发现URL中缺乏参数和片段会显著阻碍DOM-XSS检测，并证明SWIPE的DSE引擎可以合成先前未见过的URL参数和片段，从而触发20个新的漏洞。</span></span></p><p cid="n346" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1467-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1467-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n348" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">84、DUALBREACH: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization</span></span></p><p cid="n349" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">最近的研究集中在探索大型语言模型（LLMs）的漏洞，旨在从LLMs中引发有害和/或敏感内容。然而，由于针对LLMs和护栏的双重越狱攻击研究不足，当试图绕过受护栏保护的安全对齐LLMs时，现有攻击的有效性有限。因此，本文提出了DualBreach，一种面向双重越狱的目标驱动框架。DualBreach采用目标驱动初始化（TDI）策略动态构建初始提示，并结合多目标优化（MTO）方法，利用近似梯度联合调整针对护栏和LLMs的提示，从而在减少查询次数的同时实现高双重越狱成功率。对于黑盒护栏，DualBreach要么采用强大的开源护栏，要么通过训练代理模型来模拟目标黑盒护栏，从而将护栏整合到MTO过程中。通过对多个常用数据集的广泛评估，我们证明了DualBreach在双重越狱场景中的有效性。实验结果表明，DualBreach以更少的查询次数优于最先进的方法，在所有设置下都取得了显著更高的成功率。具体而言，DualBreach对受Llama-Guard-3保护的GPT-4实现了93.67%的平均双重越狱成功率，而其他方法达到的最佳成功率为88.33%。此外，DualBreach每次成功双重越狱仅使用平均1.77次查询，优于其他最先进的方法。在防御方面，我们提出了基于XGBoost的集成防御机制EGuard，该机制整合了多种护栏的优势，与Llama-Guard-3相比表现出优越的性能。</span></span></p><p cid="n350" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1062-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1062-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n352" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">85、DualStrike: Accurate, Real-time Eavesdropping and Injection of Keystrokes on Commodity Keyboards</span></span></p><p cid="n353" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们发现，在键盘上同时实现窃听和非侵入式单键注入是可行的，特别是对于快速普及的霍尔效应键盘。本文介绍了DualStrike，一种新型攻击系统，允许攻击者远程监听受害者输入并控制霍尔效应键盘上的任意按键。这种能力基于受害者的输入和上下文，开启了严重攻击（如文件删除、私钥窃取和篡改）的大门，且无需对受害者的计算机进行硬件或软件修改。我们在DualStrike中提出了几项关键创新，包括一种基于新型紧凑电磁铁的高频磁欺骗硬件设计、一种无需同步的攻击方案，以及一种使用商用现成组件的基于磁力计的监听机制。我们的真实世界实验表明，DualStrike可以可靠地攻击六种最新霍尔效应键盘模型上的任意按键。具体而言，DualStrike在所有测试模型上实现了98.9%以上的按键注入准确率。在端到端测试中，监听模块实现了高监听准确率（即超过99%）。为了提高DualStrike的鲁棒性，我们实现了一种校准算法来应对键盘位移，即使偏移达到4厘米，仍能保持98.5%的注入准确率。我们还发现了DualStrike对现有磁屏蔽机制的免疫性，并为霍尔效应键盘提出了一种新型屏蔽方法。</span></span></p><p cid="n354" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s46-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s46-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n356" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">86、Efficiently Detecting DBMS Bugs through Bottom-up Syntax-based SQL Generation</span></span></p><p cid="n357" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于语法的测试是发现数据库管理系统（DBMS）中错误的一种有前景的技术。所有现有的基于语法的SQL生成工具都采用自上而下的生成方法。为了构建SQL查询（语法树），生成器从根节点开始向前探索SQL语法，当无法将更多的语法规则应用于语法树的叶子节点时停止。然而，自上而下的生成方法倾向于投入更多精力探索接近根节点的浅层语法，而忽略了语法空间中更深层次的功能丰富的语法。因此，它在发现DBMS错误方面效率不高。本文提出了一种新的基于语法的自下而上SQL生成技术，将更多的测试资源投入到探索功能丰富的语法规则中。SQL语法的探索从一个有趣的语法规则开始，该规则概述了功能丰富的SQL功能的语法。然后，生成器将该语法规则回溯（自下而上）到根节点，创建一个揭示该有趣语法的语法路径。然后，扩展和合并多个自下而上生成的语法路径，以创建用于模糊测试的多样化SQL查询。原型工具SQLBull采用自下而上的生成技术进行模糊测试。在评估中，SQLBull在5个经过充分测试的DBMS中发现了63个零日漏洞：MySQL、MariaDB、CockroachDB、DuckDB和PostgreSQL。它在错误发现和代码覆盖率方面都优于所有现有工具。评估结果验证了自下而上生成技术的有效性。</span></span></p><p cid="n358" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f198-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f198-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n360" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">87、Enhancing Legal Document Security and Accessibility with TAF</span></span></p><p cid="n362" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">数字时代使得越来越多的服务可以通过网络访问。然而，法律获取是一个关键例外，法律仍然以纸质形式发布或发布在过时的网络平台上。采用数字法律平台的司法管辖区通常面临确保法律在线安全的困难。在本文中，我们介绍了TAF系统，该系统旨在保护法律库免受未经授权的更改，并确保法律的完整性。与以往的档案或更新框架不同，TAF是首个针对攻击者完全控制托管库这一威胁模型设计的系统。它还将每个已签名的库状态与发布者定义的法律日期绑定，从而实现可验证的特定日期检索。首先，TAF使法律文档库无论发布时间多久，都能保持可访问和可验证。其次，TAF允许任何具有库读取权限的独立验证法律库的更改。第三，TAF可供没有技术背景或网络安全知识的用户使用。TAF建立在TUF的软件更新保证、Git的版本控制结构以及强时间概念的基础上，其中时间被视为与特定库状态绑定的签名的数据。TAF将法律文档的整个演变转变为可验证、有时间戳的状态序列，确保每个过去或现在的版本都可以通过密码学方式验证。这一特性单独由Git或TUF无法提供。我们证明了TAF的安全性、可扩展性和性能，分析了其在各种攻击场景中的行为、在大法律库上的性能以及易用性。作为TAF安全性和性能的证明，TAF已被美国14个司法管辖区投入使用，包括巴尔的摩市、马里兰州和华盛顿特区。</span></span></p><p cid="n363" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1002-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1002-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n365" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">88、Enhancing Semantic-Aware Binary Diffing with High-Confidence Dynamic Instruction Alignment</span></span></p><p cid="n366" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">二进制差异检测是检测两段二进制代码之间差异的技术，是各种安全分析任务中的基础技术。现有研究表明，足够数量的细粒度对齐作为锚点可以显著提高二进制差异检测的总体准确性。然而，现有方法仍存在诸多限制，阻碍了准确高效的锚点识别。基于语法的技术容易受到激进编译优化的影响，而基于语义的方法则受限于高计算成本或低代码覆盖率。本文重新审视动态分析，寻求新的见解以解决现有方法的局限性。我们的主要见解是，并非所有动态语义对于识别有效的指令对齐都是必要或同等有效的。因此，我们可以优先使用动态执行资源，部分揭示能够有效推导指令对齐的运行时值。基于上述见解，我们提出了Barracuda，一种基于从强制执行中提取的部分指令语义的高置信度指令对齐技术。我们已实现Barracuda并进行了大量实验以评估其有效性。广泛的实验结果表明，Barracuda能够检测到24.0%更多的指令对齐作为锚点，且精度高达92.1%。Barracuda检测到的锚点可以增强最先进的二进制差异检测工具DeepBinDiff和SigmaDiff，在各种二进制差异检测场景中，F1分数分别提高了12.3%至42.7%和2.2%至4.1%。</span></span></p><p cid="n367" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f663-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f663-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n369" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">89、Enhancing Website Fingerprinting Attacks against Traffic Drift</span></span></p><p cid="n370" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">匿名通信系统，如Tor，容易受到各种网站指纹（WF）攻击的威胁，这些攻击通过分析网络流量模式来损害用户隐私。特别是，复杂的攻击采用深度学习（DL）模型来识别与特定网站相关的独特流量模式，使攻击者能够确定用户访问了哪些网站。然而，这些攻击并未设计用于处理流量漂移，如网站内容和网络条件的变化。由于流量漂移在现实生活中很常见，这些攻击在实际部署中的有效性显著降低。为解决这一局限性，我们开发了Proteus，这是第一个自适应WF攻击框架，能够在有效减轻流量漂移影响的同时，在实际场景中保持稳健的性能。Proteus的关键设计理念是仅使用漂移流量持续微调WF模型，而不需要收集部署模型时的真实标签，从而使模型能够近乎实时地适应复杂的流量漂移。具体而言，Proteus通过最小化最大均值差异来对齐原始流量和漂移流量的特征分布，并通过优化预测的熵分布来增强模型置信度。此外，它利用高斯混合模型获取可靠的伪标签，这些标签随后用于监督微调，以进一步增强其对漂移流量的鲁棒性。值得注意的是，Proteus可以与现有的基于DL的WF攻击无缝集成，以增强它们对流量漂移的适应能力。我们在包含超过35万个真实世界Tor浏览轨迹的六个流量漂移场景的大规模数据集上评估了Proteus。结果表明，对于识别漂移流量，Proteus在八种最先进的WF攻击上实现了平均94.24%的F1分数相对提升。</span></span></p><p cid="n371" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s59-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s59-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n373" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">90、Entente: Cross-silo Intrusion Detection on Network Log Graphs with Federated Learning</span></span></p><p cid="n375" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于图的入侵检测系统（GNIDS）在检测组织内部及跨边界的复杂网络攻击（如高级持续性威胁APTs）方面已展现出显著优势。尽管现有GNIDS取得了令人满意的检测精度，并能适应不断变化的攻击和正常行为模式，但它们大多假设数据集中式设置。然而，随着隐私法规约束的增加和操作限制，灵活的数据收集并不总是现实可行的。我们认为GNIDS的实际发展需要考虑分布式收集环境，并利用联邦学习（FL）作为一种可行的范式来解决这一挑战。我们观察到，将FL直接应用于GNIDS可能效果不佳，原因包括客户端图异构性以及不同GNIDS的多样化设计选择。我们提出了一系列针对图数据集的新技术来解决这些问题，包括参考图合成、图素描和自适应贡献缩放，最终开发了一个名为ENTENTE的新系统。通过利用领域知识，ENTENTE能够同时实现有效性、可扩展性和鲁棒性。在LANL、OpTC和Pivoting大规模数据集上的经验评估表明，ENTENTE优于最先进的FL基线模型。我们还评估了ENTENTE在针对GNIDS环境的FL投毒攻击下的表现，通过将攻击成功率限制在较低值，展示了其鲁棒性。总体而言，我们的研究为构建跨域GNIDS指明了一个有前景的方向。</span></span></p><p cid="n376" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s93-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s93-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n378" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">91、Eviction Notice: Reviving and Advancing Page Cache Attacks</span></span></p><p cid="n379" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">页面缓存攻击与硬件无关，并具有较高的时间和空间分辨率。自2019年以来部署的缓解措施仅保留了Evict+Reload风格的时序测量，但由于驱逐操作，这些测量方法具有极低的时间分辨率并对系统性能产生严重影响。在本文中，我们表明页面缓存攻击的问题比预期的要大得多。我们首先提出了一种基于四种基本操作的新系统化页面缓存攻击方法：刷新(flush)、重载(reload)、驱逐(evict)和监控(monitor)。基于这些基本操作，我们推导出五种针对页面缓存通用攻击技术：Flush+Monitor、Flush+Reload、Flush+Flush、Evict+Monitor和Evict+Reload。我们展示了所有基本操作的机制，这些机制可在最新的Linux内核上运行，绕过现有的缓解措施。我们在三种场景中展示了我们重新激活的页面缓存攻击的实用性，表明我们在攻击的空间和时间分辨率方面将技术水平提高了几个数量级：首先，使用我们最快的攻击方法(Flush+Monitor)，在跨进程隐蔽信道中实现了平均37.7 kB/s的信道容量。其次，对于低频攻击，我们展示了跨进程的按键间时序和事件检测攻击，空间分辨率为4 kB，时间分辨率为0.8 μs，将技术水平提高了6个数量级。第三，在网站指纹攻击中，我们在前100名的封闭世界场景中实现了90.54%的F1分数。我们得出结论，有必要针对页面缓存侧通道实施进一步的缓解措施。</span></span></p><p cid="n380" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f6-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f6-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n382" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">92、EXIA: Trusted Transitions for Enclaves via External-Input Attestation</span></span></p><p cid="n383" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可信执行环境（TEE）已被采用以保障向不可信云外包的计算安全，相关的远程认证机制使用户能够在启动时验证外包计算的完整性。然而，内存损坏攻击在启动后认证的情况下不会被检测到，从而破坏TEE的安全保证。虽然控制流认证（CFA）方案旨在检测运行时妥协，但大多数现有CFA方案缺乏具体的验证方法，且可能被仅数据攻击绕过。在本文中，我们提出了外部输入认证的概念，用于认证对TEE保护应用程序的所有写入，基于内存损坏攻击通常始于意外写入的观察。该方法通过验证所有写入符合预期来确保可信飞地状态，将控制流劫持等安全问题转化为因意外输入导致的软件崩溃等可靠性问题。为了高效地推导和验证参考测量，当前版本的外部输入认证仅限于验证者已知其输入的飞地应用程序。该设计通过在AMD SEV-SNP和Penglai上实现和评估原型得到验证，其中安全性和性能评估显示，在包括安全模型训练、模型推理、数据库工作负载和密钥管理在内的案例研究中，性能开销最小。</span></span></p><p cid="n384" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2421-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2421-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n386" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">93、Exploiting TLBs in Virtualized GPUs for Cross-VM Side-Channel Attacks</span></span></p><p cid="n387" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着虚拟GPU在云计算中的日益普及，多租户共享GPU所带来的潜在安全问题在很大程度上被忽视了。本文通过研究GPU微架构组件中的信息泄露问题，迈出了揭示这些风险的基础性一步。具体而言，我们开发了一种针对虚拟化NVIDIA GPU中后备转换缓冲区（TLBs）的Prime+Probe攻击原语。我们讨论了GPU虚拟化环境带来的几个独特挑战，并展示了我们的设计如何有效克服这些挑战。利用这一原语，我们在云环境中进行了两个跨虚拟机侧信道攻击案例研究：一个是《反恐精英2》游戏中的作弊漏洞，可以揭示隐藏的对手；另一个是网站指纹攻击，可以识别虚拟桌面用户浏览的网页。据我们所知，这些是在云环境中针对虚拟化GPU展示的首个侧信道攻击，突显了先前未知的安全风险，值得进一步研究。</span></span></p><p cid="n388" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1480-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1480-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n390" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">94、ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation</span></span></p><p cid="n391" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着大型语言模型越来越多地记忆网络爬取的训练内容，它们面临暴露版权或私人信息的风险。现有的保护措施需要爬虫或模型开发者的配合，从根本上限制了其有效性。我们提出了ExpShield，一种主动的自我保护机制，通过不可见的扰动来减轻记忆同时保持可读性，并将其表述为一个约束优化问题。由于缺乏针对自然文本的个体级风险指标，我们首先提出了实例利用（instance exploitation）这一指标，用于衡量在特定文本上进行训练会增加从一组候选文本中猜中该文本的可能性——零值表示完美的防御。对于缺乏足够知识的防御者来说，直接解决这个问题是不可行的，因此我们开发了两种有效的代理解决方案：单层优化和合成扰动。为了增强防御能力，我们揭示并验证了记忆触发假设，这有助于识别记忆的关键标记。利用这一见解，我们设计了有针对性的扰动，这些扰动（i）中和内在的触发标记以减少记忆，以及（ii）引入人工触发标记来误导模型记忆。实验验证了我们的防御在语言和视觉到语言建模中的各种攻击、模型规模和任务上的有效性。即使存在隐私后门，在防御下，成员推理攻击（MIA）的AUC值从0.95降至0.55，实例利用值接近于零。这表明，与理想的无滥用场景相比，尽管文本实例被包含在训练数据中，但其暴露的风险几乎保持不变。</span></span></p><p cid="n392" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f11-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f11-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n394" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">95、Fast Pointer Nullification for Use-After-Free Prevention</span></span></p><p cid="n395" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">像C和C++这样的低级编程语言提供了动态内存管理功能，但由于不当的释放处理，容易受到使用后释放（UAF）漏洞的攻击。这些漏洞源于通过悬空指针访问内存，构成了重大风险。尽管已经提出了各种防御机制，但现有解决方案往往面临性能开销高、内存使用过度或安全保证不足等挑战，限制了它们的实用性。指针置零（PN）作为一种有前景的UAF缓解技术，通过跟踪指针并在缓冲区释放时将其置零而受到关注。然而，现有的PN技术由于精确地将每个指针与其目标缓冲区关联而导致效率低下，造成昂贵的元数据查找。此外，它们忽略了指针存储的空间局部性，导致不必要的注册数量增加。本文介绍了快速指针置零（FPN），这是一种基于PN的新防御方法，它在区域级别组织元数据以消除昂贵的搜索操作，并使用基于块的注册来有效捕获指针局部性。在SPEC CPU基准测试和实际应用程序上的实验结果表明，与先前的PN技术相比，FPN提供了强大的安全保证，同时显著降低了性能和内存开销。FPN还兼容多线程环境和大规模Web应用程序。</span></span></p><p cid="n396" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f753-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f753-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n398" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">96、Faster Than Ever: A New Lightweight Private Set Intersection and Its Variants</span></span></p><p cid="n399" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在本工作中，我们提出了一种新的轻量级双方隐私集合交（PSI）范式，适用于半诚实模型和恶意模型。它只需要少量基础OT（不经意传输）和一次不经意键值存储（OKVS）编码和解码。所有计算（除基础OT外）均可使用SIMD加速的对称密码指令和高效的位运算实现。此外，我们将所提出的PSI协议扩展到电路PSI，并进一步扩展到多种PSI变体，包括PSI基数、PSI求和和隐私连接与计算（PJC）。所有提出的协议均在局域网（LAN）和广域网（WAN）环境下进行了评估，并与现有工作进行了性能比较。实验结果表明，在相同设置下，所提出的PSI在运行时间上比最高效的基于VOLE（不经意线性评估）的PSI快约40%，同时通信开销更低。对于电路PSI，它比基于VOLE的电路PSI构造快3.7倍，通信量减少1.5倍。在PSI基数和PSI求和的情况下，分别实现了高达12.4倍和10倍的加速，同时仅产生适度的通信开销。对于PJC，所提出的协议在运行时间上比先前工作快762倍，通信量减少3.2倍，即使在低带宽条件下也能保持高效率。</span></span></p><p cid="n400" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f131-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f131-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n402" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">97、FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation</span></span></p><p cid="n403" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">反编译是一项关键技术，可将机器码转换为人类可读格式，在没有源代码的情况下促进分析和调试。然而，这个过程面临保真度问题，这些问题会显著降低反编译输出的可读性和准确性。现有方法部分解决了这些问题，如变量重命名和结构简化，但在复杂且实用的闭源二进制场景中，通常无法提供足够的检测和纠正。为此，我们引入了FidelityGPT，这是一个新颖的框架，通过系统性地检测和纠正反编译代码与其原始源代码之间的差异，提高反编译代码的准确性和可读性。FidelityGPT定义了针对闭源环境的失真提示模板，并采用检索增强生成（RAG）技术和动态语义强度算法。该算法基于语义强度识别失真行，并从数据库中检索相似代码。此外，还设计了一种变量依赖算法，通过分析变量间的依赖关系来识别冗余变量，并将冗余变量名整合到提示上下文中，从而克服了长上下文输入的局限性。这些综合技术使FidelityGPT成为首个能够有效解决基于大语言模型的反编译优化中反编译失真问题的框架。我们在二进制相似性基准测试的620个函数对上评估了FidelityGPT，实现了89%的平均检测准确率和83%的精确率。与当前最先进的模型DeGPT（平均修复率（FR）为83%，平均修正修复率（CFR）为37%）相比，FidelityGPT表现出优越的性能，其平均FR为94%，平均CFR为64%。FidelityGPT显著提高了准确性和可读性，强调了其在增强反编译方面的有效性及其推动逆向工程发展的潜力。</span></span></p><p cid="n404" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s989-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s989-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n406" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">98、FirmAgent: Leveraging Fuzzing to Assist LLM Agents with IoT Firmware Vulnerability Discovery</span></span></p><p cid="n407" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">物联网设备的快速普及带来了严重的安全漏洞。现有的漏洞检测技术存在多种缺陷：静态分析解决方案（包括大型语言模型，LLMs）存在高误报率且无法提供概念验证(PoC)样本，而动态分析解决方案（如模糊测试）则往往存在高漏报率。为应对这些挑战，我们提出了FirmAgent，这是首个利用模糊测试辅助LLM智能体在物联网固件中查找漏洞的混合解决方案。我们的设计基于一个关键观察：模糊测试能够准确识别固件中与输入相关的代码点，而静态分析则可以彻底分析从这些代码点开始的程序路径。FirmAgent利用模糊测试收集运行时输入点（即污点源）并重建潜在的漏洞路径。然后，它应用一个LLM智能体沿着潜在路径执行上下文感知的污点分析，并应用另一个LLM智能体优化模糊测试生成的测试用例以生成概念验证测试用例。我们在14个真实物联网固件上评估了FirmAgent。它以91%的精确率识别出182个漏洞，其中包括140个先前未知的漏洞，其中17个已被分配CVE编号。我们的结果表明，FirmAgent在检测能力和精确率方面均显著优于最先进的工具。</span></span></p><p cid="n408" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1943-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1943-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n410" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">99、FirmCross: Detecting Taint-style Vulnerabilities in Modern C-Lua Hybrid Web Services of Linux-based Firmware</span></span></p><p cid="n411" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">静态污点分析已成为检测基于Linux固件Web服务中隐含漏洞的基本技术。然而，现有工作通常过于简化固件Web服务的组成。具体而言，漏洞检测范围仅考虑C二进制文件（即从目标固件中提取的二进制文件）。在本工作中，我们观察到现代固件广泛结合Lua脚本/字节码和C二进制文件来实现混合Web服务，显然，那些以C二进制文件为导向的漏洞检测技术难以取得令人满意的性能。鉴于此，我们提出了FirmCross，一个专门针对C-Lua混合Web服务的自动化污点式漏洞检测器。与现有检测器相比，FirmCross可以自动反混淆目标固件中的Lua字节码，额外识别Lua代码空间中的独特污点源，并系统性地捕获C-Lua跨语言污点流。在评估中，FirmCross在一个包含来自11个厂商的73个固件映像的数据集中，比最先进的方法（即MangoDFA和LuaTaint）多检测出6.82倍至14.5倍的漏洞。值得注意的是，FirmCross帮助在目标固件映像中识别出610个0日漏洞。在向厂商报告这些漏洞后，迄今为止已有31个漏洞ID被分配。</span></span></p><p cid="n412" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1251-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1251-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n414" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">100、FLIPPYRAM: A Large-Scale Study of Rowhammer Prevalence</span></span></p><p cid="n416" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Rowhammer是动态随机存取存储器（DRAM）中的一种扰动错误，可以通过软件故意触发，即通过反复读取（即锤击）不同DRAM行中的邻近内存位置来实现。尽管大量研究评估了Rowhammer效应，特别是其触发方式和利用方法，但大多数研究仅使用了少量双列直插式存储模块（DIMM）样本。只有少数研究提供了该效应普遍性的证据，但这些研究存在明显局限，仅限于特定硬件配置或基于FPGA的实验（这些实验能精确控制DIMM），限制了结果的泛化程度。在本文中，我们进行了首个关于Rowhammer效应的大规模研究，涉及来自822个系统的1006个数据集。我们使用最先进的基于软件的DRAM和Rowhammer工具，在一个名为FlippyRAM的全自动化跨平台框架中测量Rowhammer的普遍性。我们的框架自动收集DRAM信息，并使用5种工具来逆向工程DRAM寻址函数，然后基于这些逆向工程函数使用7种工具发起Rowhammer攻击。我们从2024年12月30日至2025年6月30日，通过在线和USB闪存驱动器向数千名参与者分发该框架。总体而言，我们从具有各种CPU、DRAM代际和供应商的系统中收集了1006个数据集。我们的研究显示，在1006个数据集中，有453个（822个独特系统中的371个）成功完成了DRAM寻址函数逆向工程的第一阶段，这表明成功且可靠地恢复DRAM寻址函数仍然是一个重大的开放性问题。在第二阶段，126个数据集（占总数据集的12.5%）在我们的全自动化Rowhammer攻击中出现了位翻转。我们的结果表明，全自动化即可武器化的Rowhammer攻击所能影响的系统比例低于基于FPGA和实验室实验所表明的比例，但12.5%的比例已足以成为威胁行为者的实用攻击向量。此外，我们的研究结果强调，围绕Rowhammer可利用性的两个最紧迫的研究挑战是：更可靠的逆向工程寻址函数（因为50%未出现位翻转的数据集在DRAM逆向工程阶段失败），以及跨多样化处理器微架构的可靠Rowhammer攻击（因为只有12.5%的数据集包含位翻转）。解决这些挑战中的每一个都可能使易受Rowhammer攻击的系统数量翻倍，并使Rowhammer在现实场景中成为更紧迫的威胁。</span></span></p><p cid="n417" mdtype="paragraph" style="box-sizing: border-box;text-align: left;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1810-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1810-paper.pdf</a></span></span></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=007d269e&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486060%26idx%3D1%26sn%3D2ed581b7ad4a96197103b393cdfea9a7">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sun, 01 Mar 2026 14:04:00 +0800</pubDate>
    </item>
    <item>
      <title>NDSS 2026论文清单及摘要（中）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486060&amp;idx=2&amp;sn=75c9796f6cfd6cf0ea4c4ffd390dd333</link>
      <description></description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span>2026-03-01 14:04</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=f03ea743&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2sz36MDumEsk7ib4ltjXjaP5M8SRaXUSMyMVFapPriacbG1Jskn3dUDfNITAF2IglC7MUiblvvraFlkOI1RT6XI7DyQk9UdJYA9f6I%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p class="mp_profile_iframe_wrp" nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="0" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><p cid="n419" mdtype="paragraph" style="box-sizing: border-box;" data-pm-slice="0 0 []"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">101、FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems</span></span></p><p cid="n420" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">自主目标跟踪（ATT）系统，尤其是ATT无人机，广泛应用于监控、边境控制和执法等领域，同时也被滥用于跟踪和破坏行为。因此，ATT的安全性对实际应用至关重要。在此背景下，我们提出了一种新型攻击：距离拉回攻击（DPA），并对其进行了系统性研究，该攻击利用ATT系统的漏洞，危险地减少跟踪距离，导致无人机被捕获、传感器攻击敏感性增加，甚至发生物理碰撞。为实现这些目标，我们提出了FlyTrap，一种新颖的物理世界攻击框架，它使用一把对抗伞作为可部署和领域特定的攻击向量。FlyTrap专门设计用于满足ATT无人机攻击的关键目标：物理可部署性、闭环有效性和时空一致性。通过新颖的渐进式距离拉回策略和可控的时空一致性设计，FlyTrap在实际环境中操控ATT无人机，实现了显著的系统级影响。我们的评估包括在真实白盒甚至商用ATT无人机（包括DJI和HoverAir）上进行的新数据集、指标和闭环实验。结果表明，FlyTrap能够将跟踪距离减少到可被捕获、传感器攻击甚至直接坠机的范围内，凸显了ATT系统安全部署的紧迫安全风险和实际意义。</span></span></p><p cid="n421" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s904-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s904-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n423" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">102、Formal Analysis of BLE Secure Connection Pairing and Revelation of the PE Confusion Attack</span></span></p><p cid="n424" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">安全连接(SC)配对是最新版本的安全协议，旨在保护通过低功耗蓝牙(BLE)信道传输的敏感信息。对该协议进行正式且严谨的分析对于提高安全保证和识别潜在漏洞至关重要。然而，协议流程的复杂性、配对方法形式化的困难以及过于理想化的用户假设为这种分析带来了重大障碍。在本文中，我们解决了这些挑战，并使用Tamarin工具对BLE-SC配对协议进行了准确且全面的正式分析。我们提取了每个参与者的状态机作为协议建模的蓝图，并使用等式理论来形式化配对方法选择逻辑。我们的模型包含了细微的用户行为，并考虑了更强的对手能力，包括对临时带外信道等私有信道的潜在妥协。我们开发了一种验证策略来自动化协议分析，并实现了一个脚本以在多个服务器上并行化验证任务。我们验证了84种配对案例，并确定了协议所需的最小安全假设。此外，我们的结果揭示了一种新的中间人(MitM)攻击，我们称之为PE混淆攻击。我们提供了在受控环境中模拟和理解此攻击的工具和概念验证(PoC)漏洞利用程序。最后，我们提出了防御此攻击的对策，提高了BLE-SC配对协议的安全性。</span></span></p><p cid="n425" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f779-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f779-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n427" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">103、From Noise to Signal: Precisely Identify Affected Packages of Known Vulnerabilities in npm Ecosystem</span></span></p><p cid="n428" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">npm是最大的开源软件生态系统，拥有超过300万个软件包。然而，包之间的复杂依赖关系使其面临严重的安全威胁，因为许多包直接或间接依赖于其他存在已知漏洞的包。及时更新这些易受攻击的依赖是软件供应链安全中的一个重大挑战，主要由于漏洞的广泛影响和修复它们的高昂成本。最近的研究表明，现有的包级漏洞传播分析工具会导致高误报率，而函数级工具在npm生态系统中尚不适用于大规模分析。在本文中，我们提出了一个新颖的框架VulTracer，它可以精确高效地执行函数级漏洞传播分析。通过为每个包独立构建丰富的语义图，然后将它们连接起来，VulTracer可以精确定位漏洞传播路径并识别真正受影响的包。通过比较评估，我们的框架在调用图构建中实现了0.905的F1分数，并将npm audit的误报率降低了94%。我们对整个npm生态系统进行了迄今为止最大规模的函数级漏洞影响测量，涵盖了3400万个包版本。结果表明，包级分析确定的68.28%的潜在影响只是噪音，因为易受攻击的代码是不可达的。此外，我们的研究还发现真正的漏洞传播（信号）是浅层的，影响在仅仅几个依赖跳转内就会显著减弱。VulTracer为缓解警报疲劳并提供了一种实用路径，使安全工作能够专注于真正可达的威胁。</span></span></p><p cid="n429" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1902-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1902-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n431" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">104、From Obfuscated to Obvious: A Comprehensive JavaScript Deobfuscation Tool for Security Analysis</span></span></p><p cid="n432" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">JavaScript的广泛应用使其成为恶意攻击者的有吸引力的目标，这些攻击者采用复杂的混淆技术来隐藏恶意代码。当前的去混淆工具存在严重局限性，严重限制了它们的实际有效性。现有工具难以处理多样化的输入格式，仅针对特定的混淆类型，并且产生晦涩难懂的输出，阻碍了人工分析。为应对这些挑战，我们提出了JSIMPLIFIER，这是一个全面的去混淆工具，采用多阶段流水线，包括预处理、基于抽象语法树的静态分析、动态执行跟踪以及大型语言模型（LLM）增强的标识符重命名。我们还引入了多维评估指标，结合了控制/数据流分析、代码简化评估、熵度量和基于LLM的可读性评估。我们构建并发布了最大的真实世界混淆JavaScript数据集，包含44,421个样本（23,212个野生恶意样本和21,209个良性样本）。评估显示，JSIMPLIFIER在处理20种混淆技术时达到100%的处理能力，在评估子集上达到100%的正确性，代码复杂度降低88.2%，并通过多个LLM验证可读性提高超过4倍。我们的成果推进了JavaScript去混淆研究和实际安全应用的基准。</span></span></p><p cid="n433" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2198-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2198-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n435" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">105、From Perception to Protection: A Developer-Centered Study of Security and Privacy Threats in Extended Reality (XR)</span></span></p><p cid="n436" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">XR的沉浸式特性引入了一组根本不同的安全与隐私（S&amp;P）挑战，这些挑战源于传统范式难以缓解的前所未有的用户交互和数据收集。作为XR应用的主要架构师，开发者在应对新型威胁方面发挥着关键作用。然而，为了有效支持开发者，我们首先必须了解他们如何感知和应对不同威胁。尽管这一问题日益重要，但缺乏从开发者角度深入考察XR安全与隐私的威胁感知研究。为填补这一空白，我们采访了23名专业XR开发者，重点关注XR中的新兴威胁。我们的研究旨在解决两个研究问题，以揭示XR开发中的现有问题并确定可行的前进路径。通过考察开发者对安全与隐私威胁的感知，我们发现：（1）XR开发决策（如丰富的传感器数据收集、用户生成内容界面）与安全与隐私威胁密切相关并可能放大这些威胁，但开发者往往没有意识到这些风险，导致威胁感知中的认知偏见；（2）现有缓解方法的局限性，加上战略、技术和沟通支持不足，削弱了开发者有效应对这些威胁的动机、意识和能力。基于这些发现，我们提出了切实可行且考虑各利益相关者的建议，以在整个XR开发过程中提升XR的安全与隐私。这项工作代表了XR领域首次进行的威胁感知、以开发者为中心的研究——XR技术的沉浸式、数据丰富特性在这一领域引入了独特的挑战。</span></span></p><p cid="n437" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s807-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s807-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n439" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">106、Fuzzilicon: A Post-Silicon Microcode-Guided x86 CPU Fuzzer</span></span></p><p cid="n440" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代中央处理器（CPU）是黑盒、专有的，并且越来越具有复杂的微架构缺陷，这些缺陷能够规避传统分析。虽然其中一些关键漏洞是通过繁琐的手动工作发现的，但为现实世界的后硅处理器构建一个自动化的系统性漏洞检测框架仍然是一个挑战。在本文中，我们提出了Fuzzilicon，这是第一个针对现实世界x86 CPU的后硅模糊测试框架，它能够深入检查微代码和微架构层。Fuzzilicon自动化了那些以前只能通过大量手动逆向工程才能发现的漏洞的检测，并通过引入微代码级检测工具弥合了可见性差距。Fuzzilicon的核心是一种从处理器微架构直接提取反馈的新技术，该技术通过逆向工程英特尔的专有微代码更新接口实现。我们开发了一种最小侵入性的检测方法，并将其基于 hypervisor 的模糊测试工具集成，以实现精确的反馈引导输入生成，无需访问寄存器传输级（RTL）或供应商支持。应用于英特尔的Goldmont微架构，Fuzzilicon发现了5个重要发现，包括两个以前未知的微代码级推测执行漏洞。此外，Fuzzilicon框架自动重新发现了先前工作中手动检测到的μSpectre类漏洞。与基线技术相比，Fuzzilicon将覆盖率收集开销降低了31倍，并实现了可挂钩位置16.27%的唯一微代码覆盖率，这是该领域的首个经验基线。作为一个实用的、覆盖率引导的、可扩展的后硅模糊测试方法，Fuzzilicon为自动化发现复杂CPU漏洞建立了新的基础。</span></span></p><p cid="n441" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1486-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1486-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n443" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">107、GoldenFuzz: Generative Golden Reference Hardware Fuzzing</span></span></p><p cid="n444" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代硬件系统因高性能和应用特定功能的需求驱动而日益复杂，引入了大量的错误和安全关键漏洞。模糊测试已成为发现此类缺陷的可扩展解决方案。然而，现有的硬件模糊测试器由于依赖缓慢的设备仿真，存在语义感知有限、测试效率低下和计算开销大等问题。本文提出了GoldenFuzz，一种新颖的两阶段硬件模糊测试框架，部分地将测试用例的精炼与覆盖率和漏洞探索解耦。GoldenFuzz利用一个快速、符合ISA规范的黄金参考模型（GRM）作为被测设备（DUT）的&#34;数字孪生&#34;。它首先对GRM进行模糊测试，实现快速、低成本的测试用例精炼，加速在DUT上的深度架构探索和漏洞发现。在模糊测试流程中，GoldenFuzz通过精心选择的指令块串联来迭代构建测试用例，这些指令块平衡了指令间和指令内的微妙质量。利用高覆盖率和低覆盖率样本见解的反馈驱动机制进一步增强了GoldenFuzz在硬件状态探索方面的能力。我们对三个RISC-V处理器（RocketChip、BOOM和CVA6）的评估表明，GoldenFuzz在实现最高覆盖率的同时，以最少的测试用例长度和计算开销显著优于现有模糊测试器。GoldenFuzz发现了所有已知漏洞和五个新漏洞，其中四个被归类为高度严重，CVSS v3严重性评分超过七分。它还识别了商业BA51-H核心扩展中的两个先前未知的漏洞。</span></span></p><p cid="n445" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1663-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1663-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n447" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">108、Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy</span></span></p><p cid="n448" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">WhatsApp作为截至2025年初拥有35亿活跃账户的平台，是全球最大的即时通讯平台。凭借庞大的用户基础，WhatsApp在全球通信中发挥着关键作用。要发起对话，用户必须首先确认其联系人是否已在该平台注册。这是通过查询WhatsApp服务器实现的，服务器会提取用户通讯录中的手机号码（如果用户已授权访问）。这种架构 inherently enables phone number enumeration，因为服务必须允许合法用户查询联系人可用性。虽然速率限制是防止滥用的标准防御措施，我们重新审视了这一问题，并表明WhatsApp在规模化枚举方面仍然存在高度漏洞。在我们的研究中，我们每小时能够探测超过一千万个电话号码而未遇到阻止或有效的速率限制。我们的研究结果不仅证明了这一漏洞的持续性，还揭示了其严重性。我们进一步发现，2021年Facebook数据泄露事件中披露的电话号码中，近一半仍然活跃在WhatsApp上，强调了此类泄露带来的持续风险。此外，我们还对WhatsApp用户进行了普查，揭示了即使消息本身是端到端加密的，大型通讯服务仍能产生的宏观洞察。利用收集的数据，我们还发现某些X25519密钥在不同设备和电话号码中被重复使用，表明存在不安全（自定义）实现或欺诈活动。</span></span></p><p cid="n449" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s805-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s805-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n451" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">109、Hiding an Ear in Plain Sight: On the Practicality and Implications of Acoustic Eavesdropping with Telecom Fiber Optic Cables</span></span></p><p cid="n452" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">光纤因其对外部干扰的抵抗能力和低信号损耗而被广泛认为是可靠的通信渠道。本文展示了电信光纤中存在的一个关键侧信道，该信道允许进行声学窃听。通过利用光纤对声振动的敏感性，攻击者可以远程监测光纤结构中由声音引起的形变，并进一步从原始声波中恢复信息。随着现代建筑中光纤到户（FTTH）安装的普及，这一问题变得尤为令人担忧。攻击者只需访问光纤的一端，即可使用商用分布式声学传感（DAS）系统窃听另一端周围的环境。然而，由于光纤本身对空气传播的声音不够敏感，我们引入了一种&#34;感官受体&#34;以提高声学捕获能力。我们的研究结果表明，能够恢复关键信息，如人类活动、室内定位和对话内容，这引发了人们对光纤通信网络隐私的重要担忧。</span></span></p><p cid="n453" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f546-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f546-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n455" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">110、HoneySat: A Network-based Satellite Honeypot Framework</span></span></p><p cid="n456" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">卫星是关键任务服务的支柱，使我们的现代社会能够正常运转，例如GPS。多年来，卫星被认为具有安全性，因为其难以理解的架构和依赖&#34;通过隐匿实现安全&#34;的策略。然而，技术进步使这些假设过时，为潜在攻击铺平了道路。不幸的是，目前无法收集有关卫星对抗技术的数据，这阻碍了导致反措施开发的情报生成。在本文中，我们提出了HoneySat，这是第一个高交互式卫星蜜罐框架，能够真实地模拟真实的立方星（CubeSat），这是一种小型卫星（SmallSat）。为了证明HoneySat的有效性，我们调查了小型卫星运营商并通过互联网部署了HoneySat。我们的研究结果显示，90%的卫星运营商同意HoneySat提供了真实的模拟。此外，HoneySat成功欺骗了现实世界中的攻击者，并收集了22个真实的对抗性交互。最后，我们进行了硬件在环操作，其中HoneySat成功与在轨运行的小型卫星任务进行了通信。</span></span></p><p cid="n457" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f537-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f537-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n459" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">111、HOUSTON: Real-Time Anomaly Detection of Attacks against Ethereum DeFi Protocols</span></span></p><p cid="n460" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着去中心化金融（DeFi）持续创新金融体系，其基础构件的安全性仍是其大规模应用的关键关注点。在DeFi领域，风险极高，每周都有数百万美元的财务损失事件反复发生。所有主要的基于区块链的金融应用（即DeFi协议）都由称为智能合约的程序构建并与之交互。虽然已开发了许多安全工具来识别单个智能合约中的特定漏洞类别（如重入攻击），但在自动实时识别针对DeFi协议的攻击方面，投入的努力相对较少。在本文中，我们提出了一种新颖的方法，用于实时、通用且可解释地识别针对DeFi协议的攻击。具体而言，我们识别潜在的风险交易，而不依赖于任何已知的漏洞模式。我们的方法在HOUSTON系统中实现，首先自动识别共同实现DeFi应用的智能合约集合，然后在监控新的相关交易时，构建和更新自定义异常检测模型。我们的模型包含典型执行路径（控制流）的信息，以及协议如何处理数据的信息，这些信息被捕获为合约函数参数与存储变量之间可能的不变量关系。HOUSTON提供可解释的警报，可用于攻击分类。我们在超过2200万笔交易的大型语料库上评估了HOUSTON，涵盖了115个DeFi事件。在我们的实验中，HOUSTON实现了94.8%的检测真阳性率，同时保持低假阳性率。与最先进的异常检测系统相比，HOUSTON实现了更高的真阳性数量和更低的假阳性率。最后，我们在真实环境中部署了HOUSTON，它在普通硬件上展示了实时监控能力，同时保持了高准确性。</span></span></p><p cid="n461" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1534-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1534-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n463" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">112、Huma: Censorship Circumvention via Web Protocol Tunneling with Deferred Traffic Replacement</span></span></p><p cid="n464" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着互联网审查日益普遍，用户常常依赖隐蔽通道来规避监控并访问受限内容。Web协议隧道工具使用网站作为代理，将隐蔽数据封装在Web协议中，以与合法流量混合从而避免检测。然而，现有工具容易通过流量分析被检测到，使审查者能够通过指纹攻击或因产生异常浏览模式来识别此类工具的使用。我们提出了Huma，一种新的Web协议隧道工具，解决了现有的检测问题。通过延迟隐蔽数据传输，Huma允许参与规避审查的网站首先返回未修改的内容，而嵌入隐蔽数据的响应则在后台准备并在客户端的下一个请求期间发送，从而避免了促进指纹识别的时间异常。通过依赖基于真实浏览活动建模的显式用户模拟器，Huma也遵循用户预期的浏览行为。最后，Huma防止对手控制的网站将通信端点绑定在一起，从而能够轻松扩展以支持内部网审查场景中的隐蔽通信。</span></span></p><p cid="n465" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f328-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f328-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n467" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">113、HyperMirage: Direct State Manipulation in Hybrid Virtual CPU Fuzzing</span></span></p><p cid="n468" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">虚拟机监视器对现代云基础设施的安全性和可用性至关重要，但它们必须向客户虚拟机暴露大量的虚拟化接口——这是攻击者可以利用的攻击面。虚拟机监视器中最复杂且对安全敏感的组件之一是其虚拟CPU实现，通常在最高特权级别实现。尽管之前的模糊测试研究在检查虚拟机监视器的虚拟CPU组件方面取得了有希望的进展，但现有技术无法深入覆盖该组件，因为其复杂的性质需要繁琐的手动设置来访问各个接口，同时采用次优技术降低了模糊测试吞吐量。我们通过HyperMirage解决了这些缺陷，这是一种新型虚拟机监视器模糊测试工具，能够自动高效地探索虚拟CPU实现所模拟的大量架构状态空间。HyperMirage采用一种新颖的直接状态操作方法，使安全分析师无需手动构建架构有效的虚拟机状态作为模糊测试种子，该方法直接且自动地修改虚拟机监视器在模糊测试过程中所使用的虚拟机状态视图。此外，我们扩展了最先进的基于编译器的符号执行引擎，使其成为首个可用于裸机目标的引擎，并将其集成到高效的覆盖率引导虚拟机监视器模糊测试工具中，使HyperMirage与现有技术相比能够显著提高模糊测试吞吐量。我们通过在Intel x86架构上对生产级Xen和KVM虚拟机监视器进行模糊测试，提供了HyperMirage的案例研究。我们的评估表明，HyperMirage能够比先前工作多覆盖200%的虚拟CPU接口，与可用的虚拟机监视器模糊测试工具相比，在整个虚拟CPU空间上实现了显著更高的覆盖率。此外，HyperMirage在Xen中发现了9个新漏洞，在KVM中发现了2个新漏洞，所有这些漏洞都已被各自的项目维护者确认。</span></span></p><p cid="n469" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1763-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1763-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n471" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">114、Icarus: Achieving Performant Asynchronous BFT with Only Optimistic Paths</span></span></p><p cid="n472" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">区块链技术的出现重新激发了人们对拜占庭容错（BFT）共识的研究兴趣，特别是异步BFT，因为它对网络攻击具有抵抗力。为了提高传统异步BFT的性能，最近的研究提出了双路径范式：在有利情况下通过乐观路径提高效率，在不利情况下通过悲观路径（通常通过多值验证拜占庭协议（MVBA）实现）保证活性。然而，由于MVBA协议固有的复杂性和低效性，现有的双路径协议在不利情况下表现出高实现复杂性和性能差。此外，双路径范式中的两种构成类型——串行路径和并行路径——各自面临额外的限制。具体而言，串行类型在乐观路径和悲观路径之间切换困难，而并行类型会丢弃其中一个路径的区块，导致带宽浪费和吞吐量降低。为解决这些限制，我们提出了Icarus，这是一种单路径异步BFT协议，仅利用乐观路径而不使用悲观路径。乐观路径确保Icarus在有利情况下的效率。为保证不利条件下的活性，Icarus采用旋转链机制：每个节点并行广播一个区块链，这些链以轮询方式轮流作为乐观路径。由于无故障节点的链持续增长，一旦积累了足够区块的链成为乐观路径，其区块就可以被提交，从而确保即使在不利条件下也能保持活性。为在路径转换过程中保持一致性，Icarus引入了双连续验证值拜占庭协议（tcv$^2$-BA），该协议对先前路径上已提交区块的高度进行对齐。我们通过理论分析验证了Icarus的正确性，并通过各种实验证明了其高性能。</span></span></p><p cid="n473" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f60-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f60-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n475" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">115、Identifying Logical Vulnerabilities in QUIC Implementations</span></span></p><p cid="n476" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">QUIC是一种现代传输协议，正被各大平台和服务越来越多地采用，因此其安全性和正确性至关重要。然而，QUIC规范和实现的复杂性引入了细微且危险逻辑缺陷的机会。现有的QUIC测试工具主要关注与内存相关的漏洞，而难以检测逻辑漏洞。因此，逻辑漏洞的发现目前仍然高度依赖人工审计。在本文中，我们介绍了MerCuriuzz，这是一种新颖的黑盒模糊测试框架，旨在自动发现QUIC实现中的逻辑漏洞。我们对16种广泛使用的QUIC实现进行了MerCuriuzz评估，发现了14个先前未知的逻辑漏洞，这些漏洞影响了quiche、xquic和aioquic等流行实现。这些漏洞可能带来严重的安全风险，使攻击者能够耗尽服务器资源、使服务崩溃或拒绝合法用户访问服务器。我们将这些漏洞分为六类，并提出了缓解策略。我们还负责任地向相关供应商披露了我们的发现，其中11个漏洞已被供应商确认并获得奖励，例如Cloudflare和阿里云。</span></span></p><p cid="n477" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1777-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1777-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n479" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">116、Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined Types</span></span></p><p cid="n480" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">反编译器帮助逆向工程师在比汇编代码更高抽象层次上分析软件。不幸的是，由于编译过程会丢失信息，传统的确定性反编译器生成的代码缺乏许多使源代码具有可读性的特性，例如变量和类型名称。神经反编译器提供了通过统计方法填补这些细节的可能性。然而，现有的神经反编译工作存在重大局限，使其无法应用于真实代码，例如无法为用户定义的复合类型提供定义。在这项工作中，我们介绍了Idioms，这是一种简单、可推广且有效的神经反编译方法，可以将任何大型语言模型微调为能够生成适当用户定义类型定义以及反编译代码的神经反编译器，同时我们还创建了一个新数据集Realtype，其中包含比现有神经反编译基准测试更复杂和更真实的类型。我们证明，我们的方法在神经反编译领域取得了最先进的结果。在最具挑战性的现有基准测试Exebench上，我们的模型达到了54.4%的准确率，而LLM4Decompile为46.3%，Nova为37.5%；在Realtype上，我们的模型性能提升了至少95%。</span></span></p><p cid="n481" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f795-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f795-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n483" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">117、In-Context Probing for Membership Inference in Fine-Tuned Language Models</span></span></p><p cid="n484" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">成员推断攻击（MIAs）对微调的大型语言模型（LLMs）构成了严重的隐私威胁，特别是当模型使用敏感数据针对特定领域任务进行适配时。虽然先前的黑盒MIA技术依赖于置信度分数或令牌似然度，但这些信号通常与样本的固有特性（如内容难度或稀有性）相互纠缠，导致泛化能力差且信噪比低。在本文中，我们提出了ICP-MIA，这是一个基于训练理论的新颖MIA框架，特别关注优化过程中出现的收益递减现象。我们引入了优化差距作为成员的基本信号：在收敛时，成员样本表现出最小的剩余损失降低潜力，而非成员则保留显著的进一步优化潜力。为了在黑盒环境中估计这一差距，我们提出了上下文探测（ICP）——一种无需训练的方法，通过 strategically 构建的输入上下文模拟类似微调的行为。我们提出了两种探测策略：基于参考数据（使用语义相似公共样本）和自扰动（通过掩码或生成）。在三个任务和多个LLMs上的实验表明，ICP-MIA显著优于先前的黑盒MIA，特别是在低误报率的情况下。我们进一步分析了参考数据对齐、模型类型、PEFT配置和训练计划如何影响攻击效果。我们的研究结果表明，ICP-MIA是一个实用的、有理论基础的框架，可用于评估已部署LLMs的隐私风险。</span></span></p><p cid="n485" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f892-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f892-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n487" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">118、Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination</span></span></p><p cid="n488" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">及时有效的应急响应是应对日益增多的网络攻击的关键。然而，为复杂系统确定正确的响应措施是一项重大技术挑战。缓解这一挑战的一种有前景的方法是利用嵌入大型语言模型（LLM）中的安全知识来协助安全操作员在事件处理过程中的工作。最近的研究已经证明了这种方法的可能性，但当前的方法主要基于前沿LLM的提示工程，这种方法成本高昂且容易出现幻觉。我们通过提出一种使用LLM进行应急响应规划的新方法来减少幻觉，从而解决这些局限性。我们的方法包括三个步骤：微调、信息检索和前瞻性规划。我们证明，在特定假设条件下，我们的方法生成的响应计划具有有限的幻觉概率，并且可以通过增加规划时间使这种概率任意小。此外，我们展示了我们的方法是轻量级的，可以在普通硬件上运行。我们在文献中报道的事件日志上评估了我们的方法。实验结果表明，我们的方法a）比前沿LLM缩短高达22%的恢复时间，并且b）能够广泛适用于各种事件类型和响应措施。</span></span></p><p cid="n489" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f358-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f358-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n491" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">119、Indicator of Benignity: An Industry View of False Positive in Malicious Domain Detection and its Mitigation</span></span></p><p cid="n492" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">恶意域名检测是保护用户免受网络攻击的关键技术。尽管这些系统已展现出显著的检测能力，但它们在现实世界中的误报（FPs）规模仍然未知，且常被忽视。为了阐明这一重要方面，我们进行了一项首次测量研究，使用了从全球最大的网络安全供应商之一收集的6年误报报告。我们的研究结果表明，当前检测系统普遍采用的基于流行度的顶级域名列表不足以避免误报。事实上，在生产环境中仍存在大量误报。我们认为，主要原因之一是该领域的努力主要集中在检测恶意指标（即入侵指标，IOC）上，而忽视了良性指标（即良性指标，IOB）。在本文中，我们首次专注于IOB检测的研究。我们的工作基于一个关键发现：对于生产环境中的许多误报，其IOB可以在互联网上找到。然而，由于互联网的开放性和网络内容的不结构化，我们在识别这些IOB时面临两个主要挑战：理解IOB是什么以及评估IOB的可信度。为应对这些挑战，我们提出了一个IOB的传递信任模型，并在名为IOBHunter的系统中实现了该模型。IOBHunter利用了大语言模型（LLM）和思维链（CoT）技术，这些技术已在解决其他几种安全威胁方面展现出良好的能力。我们使用包含已验证误报的数据集进行的评估显示，IOBHunter可以达到99.22%的精确率和68.6%的召回率。IOBHunter还在为期两个月的实际部署中进行了进一步评估，期间IOBHunter识别出了4,338个已确认的误报和2,051个被攻陷的域名。</span></span></p><p cid="n493" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1869-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1869-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n495" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">120、InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation Analysis</span></span></p><p cid="n496" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">像CLIP这样的多模态对比学习模型展示了卓越的视觉-语言对齐能力，现已成为许多大规模多模态系统的基础组件。然而，它们对后门攻击的脆弱性带来了严重的安全风险。攻击者可以植入潜伏的触发器，这些触发器能在下游任务中持续存在，从而在触发器出现时实现对模型行为的恶意控制。尽管最近的防御机制取得了巨大成功，但由于对攻击者知识的强假设或对干净数据的过度需求，它们仍然不切实际。在本文中，我们提出了InverTune，这是首个在最小攻击者假设条件下的多模态模型后门防御框架，既不需要攻击目标的先验知识，也不需要访问被污染的数据集。与依赖于中毒阶段使用的数据集的现有防御方法不同，InverTune通过三个关键组件有效识别并移除后门工件，从而实现对后门攻击的强大保护。具体而言，（1）InverTune首先通过对抗性模拟暴露攻击特征，通过分析模型响应模式概率性地识别目标标签。（2）在此基础上，我们开发了一种梯度反转技术，通过激活模式分析来重建潜伏的触发器。（3）最后，采用聚类引导的微调策略，仅使用少量任意干净数据来消除后门功能，同时保留原始模型能力。实验结果表明，InverTune将最先进（SOTA）攻击的平均攻击成功率（ASR）降低了97.87%，同时将干净准确率（CA）的 degradation限制在仅3.07%。这项工作为保障多模态系统安全建立了新范式，推进了基础模型部署的安全性，同时不损害性能。</span></span></p><p cid="n497" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1666-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1666-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n499" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">121、IoTBec: An Accurate and Efficient Recurring Vulnerability Detection Framework for Black Box IoT devices</span></span></p><p cid="n500" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">物联网设备的激增导致了漏洞利用的增加。现有的漏洞检测方法严重依赖固件或源代码进行分析，这种依赖严重限制了它们在实际黑盒场景中的效率。为解决这一局限性，我们提出了IoTBec，一种新颖的、不依赖固件和源代码的循环漏洞检测框架。IoTBec创新性地基于黑盒接口和已知漏洞信息构建了漏洞接口签名（VIS），该签名用于将潜在的循环漏洞与目标设备进行匹配。该框架随后将基于签名的检测与大型语言模型（LLM）驱动的模糊测试深度融合。当匹配成功时，IoTBec自动利用LLMs生成针对性的模糊测试载荷进行验证。为评估IoTBec，我们在来自五大物联网厂商的设备上进行了广泛实验。结果表明，IoTBec发现的漏洞数量比当前最先进的（SOTA）黑盒模糊测试方法多7倍以上，精确度为100%，召回率为93.37%。总体而言，IoTBec检测到183个漏洞，其中169个被分配了CVE ID。在这些漏洞中，53个是新发现的，平均CVSS 3.x评分为8.61，涵盖了缓冲区溢出、命令注入和CSRF问题。值得注意的是，通过LLM驱动的模糊测试，IoTBec还发现了25个先前未知的漏洞。实验证据表明，IoTBec独特的固件和源代码独立范式提高了检测效率，并能够发现新型和变体漏洞。我们将在</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/IoTBec" target="_blank">https://github.com/IoTBec</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">上发布IoTBec的源代码和实验数据。</span></span></p><p cid="n501" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f634-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f634-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n503" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">122、Ipotane: Balancing the Good and Bad Cases of Asynchronous BFT</span></span></p><p cid="n504" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">最先进的异步拜占庭容错（BFT）协议集成了部分同步的乐观路径。其最终目标是在有利情况下匹配部分同步协议的性能，在不利情况下匹配纯异步协议的性能。尽管先前的研究在有利情况下表现出色，但在条件不利时却存在不足。为解决这些缺点，最近的一项工作Abraxas（CCS&#39;23）在所有情况下都保持了稳定的吞吐量，但由于乐观路径故障检测缓慢，在不利情况下造成了极高的最坏情况延迟。另一项最近的工作ParBFT（CCS&#39;23）确保了所有情况下的良好延迟，但由于使用了额外的异步二进制协议（ABA）实例，在不利情况下吞吐量降低。我们提出了Ipotane协议，在吞吐量和延迟两方面，在有利情况下实现了与部分同步协议相当的性能，在不利情况下实现了与纯异步协议相当的性能。Ipotane同时运行两条路径：2-chain HotStuff作为乐观路径，以及一个新的原始双功能拜占庭协议（DBA）作为悲观路径。DBA封装了有偏ABA和验证异步拜占庭协议（VABA）的功能。在Ipotane中，如果副本的乐观路径更快，则向DBA输入0；如果悲观路径更快，则输入1。DBA的ABA功能通过输出1及时发出乐观路径故障的信号，确保Ipotane在不利情况下的低延迟。同时，Ipotane执行DBA实例，通过其VABA功能持续产生悲观区块。在检测到故障时，Ipotane提交最后两个悲观区块以保持高吞吐量。此外，Ipotane利用DBA的有偏特性来确保提交悲观区块的安全性。大量实验验证了Ipotane在所有情况下的高吞吐量和低延迟。</span></span></p><p cid="n505" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s3-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s3-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n507" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">123、IsolatOS: Detecting Double Fetch Bugs in COTS RTOS by Re-enabling Kernel Isolation</span></span></p><p cid="n508" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">双重获取漏洞是指内核反复从用户空间内存中检索数据，而未确保连续数据获取之间的一致性。这一问题在实时操作系统（RTOS）中尤为严重，因为严格的时序要求限制了互斥锁等同步机制的使用，从而倾向于以牺牲安全性为代价实现低延迟内存访问。大多数当前检测技术采用静态源代码分析，无法应用于具有专有内核的商业现成（COTS）RTOS。因此，采用启发式时间窗口阈值来检测跨边界内存重复访问的动态方法被采用。然而，这些方法由于模式识别过于宽泛，常常产生大量误报，并导致显著的仿真开销。我们引入了IsolatOS，一种硬件支持的检测方法，利用内核隔离功能来指示双重获取漏洞的跨边界内存访问。主要难点在于在不导致RTOS系统崩溃的情况下强制执行隔离边界的同时保持透明度，以提高效率。IsolatOS首先通过实现动态仪器来拦截对用户内存的特权访问，记录访问的元数据，然后通过异常恢复技术在故障处理期间维持系统稳定性。在执行后阶段，因果分析检查违规轨迹，以区分合法的双重访问和可利用的双重获取。在QNX、VxWorks和seL4上的评估证明了IsolatOS的有效性，与基于仿真的方法相比，运行时开销降低了70倍，识别出42个独特漏洞（39个供应商确认，2个分配的CVE）。这些结果验证了硬件辅助的内核隔离是COTS RTOS环境中双重获取检测的可行范式。我们还通过利用这些发现展示了其在汽车系统中的实际影响。</span></span></p><p cid="n509" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s568-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s568-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n511" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">124、Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds</span></span></p><p cid="n513" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">RDMA云日益普及，访问控制列表（ACL）对于规范RDMA应用、服务和租户的未授权网络访问至关重要。然而，RDMA独特的队列对（QP）语义和高传输特性使得现有的ACL表达式和执行机制无法以用户友好的方式全面高效地管理RDMA流量。在本文中，我们提出了Janus，一个专为RDMA云设计的定制化ACL系统。Janus设计了具有QP语义的专用ACL表达式来识别RDMA连接，并提供了一种高级策略语言用于表达复杂的ACL意图以管理RDMA流量。Janus进一步利用具有流量感知和架构特定优化的DPU来执行ACL策略，实现了线速RDMA检查和稳健的策略更新。我们使用NVIDIA BlueField-3 DPU实现了Janus的开源原型。实验表明，Janus为管理未授权RDMA访问提供了足够的表达能力，并在200Gbps真实RDMA测试环境中实现了线吞吐量且延迟小于5µs。</span></span></p><p cid="n514" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f721-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f721-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n516" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">125、Kangaroo: A Private and Amortized Inference Framework over WAN for Large-Scale Decision Tree Evaluation</span></span></p><p cid="n517" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着模型即服务（Models-as-a-Service）的快速采用，数据和模型隐私问题变得越来越关键。为解决这些问题，各种隐私保护推理方案已被提出。特别是由于决策树的高效性和可解释性，私有决策树评估（PDTE）已引起广泛关注。然而，现有的PDTE方案存在显著局限性：其通信和计算成本随树的数量、节点数或树深度而扩展，这使得它们对于大规模模型（尤其是在广域网环境中）效率低下。为解决这些问题，我们提出了Kangaroo，这是一个基于打包同态加密的私有且分摊的决策树推理框架。具体而言，我们设计了一种新颖的模型隐藏和编码方案，结合安全特征选择、 oblivious 比较和安全路径评估协议，实现了随着节点数或树数量增加时开销的完全分摊。此外，我们通过优化（包括相同模型共享、延迟感知和自适应编码调整策略）提升了框架的性能和功能。在广域网环境中，Kangaroo比最先进的一次性交互方案实现了14倍至59倍的性能提升。对于大规模决策树推理任务，与现有方案相比，它实现了3倍至44倍的加速。值得注意的是，在广域网环境下，Kangaroo能够以每树约60毫秒（分摊）的速度评估包含969棵树和411,825个节点的随机森林。</span></span></p><p cid="n518" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s892-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s892-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n520" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">126、Know Me by My Pulse: Toward Practical Continuous Authentication on Wearable Devices via Wrist-Worn PPG</span></span></p><p cid="n521" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">利用生理信号进行生物特征认证为可穿戴设备的安全且用户友好的访问控制提供了有前景的途径。虽然心电图（ECG）信号已显示出高度的可区分性，但其侵入式传感要求和间断性采集限制了其实用性。另一方面，光电容积脉搏波描记法（PPG）能够实现连续、非侵入式的认证，并可无缝集成到手腕可穿戴设备中。然而，大多数先前的研究依赖于高频PPG（例如75-500赫兹）和复杂的深度模型，这会导致显著的能耗和计算开销，阻碍了其在功率受限的实际系统中的部署。在本文中，我们首次在智能手表We-Be Band上实现了连续认证系统的实际部署和评估，该系统使用低频（25赫兹）多通道PPG信号。我们的方法采用带有注意力机制的Bi-LSTM从4通道PPG的短时（4秒）窗口中提取身份特定特征。通过对公共数据集（PTTPPG）和我们自己的We-Be数据集（26名受试者）的广泛评估，我们展示了强大的分类性能，平均测试准确率为88.11%，宏F1得分为0.88，错误接受率（FAR）为0.48%，错误拒绝率（FRR）为11.77%，等错误率（EER）为2.76%。与512赫兹和128赫兹的设置相比，我们的25赫兹系统将传感器功耗分别降低了53%和19%，同时不牺牲性能。我们发现25赫兹的采样率保持了认证准确性，而20赫兹时性能急剧下降，仅提供微不足道的额外节能，这凸显了25赫兹作为实际下限的重要性。此外，我们发现仅使用静息数据训练的模型在运动状态下表现不佳，而活动多样化的训练则提高了在不同生理状态下的鲁棒性。</span></span></p><p cid="n522" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1087-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1087-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n524" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">127、KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance Analysis</span></span></p><p cid="n525" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">网络威胁情报（CTI）报告中的高级自然语言知识，如ATT&amp;CK框架，有助于应对高级持续性威胁（APT）攻击。然而，如何在现实世界的攻击检测系统中（如溯源分析系统）自动应用CTI报告中的高级知识，仍然是一个开放性问题。这一挑战源于知识与低级安全日志之间的语义差距：CTI报告中的知识以自然语言形式编写，而攻击检测系统只能处理文件访问或网络IP操作等低级系统事件。手动方法可能劳动密集且容易出错。在本文中，我们提出了KnowHow，一种由CTI知识驱动的在线溯源分析方法，可以自动将CTI报告中以自然语言编写的高级攻击知识应用于检测低级系统事件。KnowHow的核心是一种新颖的攻击知识表示方法——通用入侵指标（gIoC），它表示攻击的主体、客体和行动。通过将系统事件中的系统标识符（如文件路径）提升为自然语言术语，KnowHow可以将系统事件与gIoC匹配，并进一步将其与以自然语言描述的技术进行匹配。最后，基于与系统事件匹配的技术，KnowHow对攻击步骤的时间逻辑进行推理，并在系统事件中检测潜在的APT攻击。我们的评估表明，KnowHow能够准确检测开源数据集和工业数据集中的所有16个APT活动，而现有方法都引入了大量误报。同时，我们的评估也表明，KnowHow最多减少了90%的节点级误报，同时具有更高的节点级召回率，并且对几种未知攻击和模仿攻击具有鲁棒性。</span></span></p><p cid="n526" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s199-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s199-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n528" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">128、LatticeBox: A Hardware-Software Co-Designed Framework for Scalable and Low-Latency Compartmentalization</span></span></p><p cid="n530" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代软件系统日益依赖隔离机制来隔离不可信或潜在脆弱的组件，如第三方驱动程序和即时编译代码。然而，现有的硬件隔离技术面临可扩展性限制、高切换延迟和安全性不足等问题。特别是，某些隔离技术使用的权限更改指令（如Intel MPK的WRPKRU）可能被不可信代码利用，从而增加了安全部署的复杂性。在本文中，我们介绍了LatticeBox，这是一个基于硬件-软件协同设计的框架，采用基于格的访问控制模型来解决这些局限性。LatticeBox将权限和内存区域编码为紧凑的分层N位向量。这种设计实现了硬件架构，将域切换延迟降低到单个CPU周期，并从根本上防止了权限切换指令的滥用。此外，LatticeBox采用定制指令（lp_land）来强制严格的跨域控制流完整性，有效防止未授权的间接函数调用。我们在RISC-V BOOM核心上实现了LatticeBox，并使用微基准测试和实际应用程序（包括WebAssembly运行时和Linux内核模块）对其进行了评估。结果表明，LatticeBox的域切换速度比Intel MPK快达180倍，同时支持细粒度、可扩展的隔离。在实际工作负载上的评估显示性能影响较小，增强的WebAssembly运行时仅降低2%的性能，而运行隔离Linux内核模块的ApacheBench吞吐量仅降低3%。</span></span></p><p cid="n531" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f515-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f515-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n533" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">129、Learning from Leakage: Database Reconstruction from Just a Few Multidimensional Range Queries</span></span></p><p cid="n534" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可搜索加密(SE)在实现加密数据的安全高效查询方面展现出巨大潜力。为实现这种效率，SE不可避免地会泄露一些信息，而一个重大的开放性问题在于这种泄露的危险程度如何。尽管先前的重构攻击在一维范围查询设置中已显示出有效性，但将其扩展到高维数据集仍面临挑战。现有方法要么需要过多的查询信息（例如，观察到所有可能响应的攻击者），要么在稀疏数据库中产生低质量的重构。在这项工作中，我们提出了REMIN，一种针对多维设置中SE方案的新型滥用泄露攻击，利用范围查询中的访问和搜索模式泄露。REMIN利用无监督表示学习将查询共现频率转换为几何信号，使攻击者能够推断加密记录之间的相对空间关系。这种方法在最小泄露的情况下实现了高维数据集的准确且可扩展的重构。此外，我们引入了REMIN-P，这是一种包含实用 poisoning 策略的攻击主动变体。通过注入少量辅助锚点，REMIN-P显著提高了重构质量，特别是在数据空间的稀疏或边界区域。我们在合成和真实数据集上对我们的攻击进行了广泛评估。与最先进的重构攻击相比，我们的重构攻击将均方误差(MSE)降低了高达50%，同时保持了快速且可扩展的运行时间。根据 poisoning 策略的不同，我们的 poisoning 攻击可以进一步将平均MSE额外降低50%。</span></span></p><p cid="n535" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f935-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f935-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n537" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">130、Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents</span></span></p><p cid="n538" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLM）代理是由LLM驱动的自主系统，能够利用一系列工具进行推理和规划以解决问题。然而，在LLM代理中集成多种工具带来了安全管理的挑战，包括确保工具兼容性、处理依赖关系以及保护LLM代理任务工作流中的控制流。在本文中，我们首次对多工具支持的LLM代理中的任务控制流进行了系统性的安全分析。我们识别出一种新型威胁——跨工具收集与污染（XTHP），该威胁包含多种攻击向量，首先劫持代理任务的正常控制流，然后收集并污染LLM代理系统中的机密或私有信息。为理解此威胁的影响，我们开发了Chord，一个动态扫描工具，旨在自动检测易受XTHP攻击的现实世界代理工具。我们对来自两大LLM代理开发框架LangChain和LlamaIndex的66个现实世界工具的评估显示，75%的工具易受XTHP攻击攻击，凸显了该威胁的普遍性。</span></span></p><p cid="n539" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f577-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f577-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n541" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">131、Light into Darkness: Demystifying Profit Strategies Throughout the MEV Bot Lifecycle</span></span></p><p cid="n542" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">由于无许可区块链的透明性，机会主义交易者可以通过竞争盈利机会并创建MEV机器人来使这一过程永不停歇，从而提取最大可提取价值(MEV)。然而，这种行为损害了区块链系统的共识安全性和效率。因此，了解MEV机器人的行为策略对于防范其危害至关重要。不幸的是，现有工作主要集中在MEV市场的宏观测量上，而MEV机器人策略的具体类型和分布仍然未知。在本文中，我们开发了APOLLO工具，用于分析机器人整个生命周期中的细粒度策略，从而首次对MEV机器人盈利策略进行了系统性研究。我们对2,052个MEV机器人的大规模分析得出了许多新的见解。特别是，我们首次介绍了野外机器人使用的20种代码级策略，在智能合约反混淆方面迈出了第一步，以发现隐藏在混淆机器人代码中的策略，并发现了五种能为MEV机器人带来盈利机会的特定类型交易。</span></span></p><p cid="n543" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s506-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s506-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n545" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">132、Light2Lie: Detecting Deepfake Images Using Physical Reflectance Laws</span></span></p><p cid="n546" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">生成模型（如GAN和基于扩散的架构）的快速发展导致了超写实合成图像的广泛创建。尽管这些技术推动了媒体和数据生成的创新，但也引发了重大的伦理、社会和安全问题。对此，已开发出多种检测方法，包括频域分析和深度学习分类器。然而，这些方法通常难以推广到未见过的生成模型，且往往缺乏物理基础，使其容易受到自适应攻击的影响，并且在可解释性方面存在局限。我们提出了Light2Lie，这是一个物理增强的深度伪造检测框架，它利用镜面反射原理，特别是菲涅耳反射率模型，来揭示生成模型难以有效重现的光-表面相互作用中的不一致性。我们的方法首先采用神经网络估计表面基础反射率，然后导出一种受微面启发的镜面响应图，该图编码了真实图像与合成图像之间微妙的几何和光学差异。该信号作为特征图被整合到二级分类器中，使其能够学习基于反射率驱动模式来区分两类图像。为进一步增强鲁棒性，我们引入了一种反馈细化机制，利用分类错误更新基础反射率模型的输出，将物理建模与学习目标紧密耦合。在多个深度伪造数据集上的广泛实验表明，我们的方法在处理未见过的生成模型样本时获得了更好的泛化性能，在多样化的深度伪造领域达到高达74%的精确率，优于最先进的基线方法，同时提供稳健的、基于物理的决策。</span></span></p><p cid="n547" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s923-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s923-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n549" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">133、Lightening the Load: A Cluster-Based Framework for A Lower-Overhead, Provable Website Fingerprinting Defense</span></span></p><p cid="n550" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">网站指纹识别（WF）攻击仍然是加密流量的一大威胁，促使开发了广泛的防御措施。其中，两类突出的防御方法是基于正则化的防御，它使用固定的填充规则来塑造流量，以及基于超序列的方法，它将痕迹隐藏在预定义的模式中。在这项工作中，我们提出了一个统一的框架，用于设计自适应WF防御，该方法结合了正则化的有效性和超序列式分组的可证明安全性。该方案首先从痕迹中提取行为模式，并将它们聚类到$(k,l)$-多样匿名集中；然后，一个早期时间序列分类器（从ECDIRE改编）从保守的全局正则化参数集切换到更轻量级的特定参数集。我们将该设计实例化为自适应塔马劳（Adaptive Tamaraw），它是Tamaraw的一个变体，在保留其原始信息论保证的同时，按聚类分配填充参数。在公共真实世界数据集上的全面实验证实了其优势。通过调整$k$，操作者可以在隐私和效率之间进行权衡：在其高隐私模式下，自适应塔马劳将任何攻击者的准确率上限推至低于30%，而在以效率为中心的设置中，与经典塔马劳相比，它将总开销减少了99个百分点。</span></span></p><p cid="n551" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1760-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1760-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n553" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">134、Lightweight Internet Bandwidth Allocation and Isolation with Fractional Fair Shares</span></span></p><p cid="n555" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">确保公共互联网上的公平带宽分配具有挑战性。拥塞控制算法(CCA)通常无法实现公平性，特别是当不同CCA同时运行时。在分布式拒绝服务(DDoS)攻击期间，这一挑战变得更加突出，合法流量可能完全被饿死。解决这一挑战的一种方法是通过在路由器上直接分配带宽来强制执行公平性。然而，现有解决方案通常分为两类：一类易于部署但无法提供安全的网络内带宽隔离，另一类提供强大的隔离保证但依赖于阻碍实际部署的复杂假设。为了弥合这两类解决方案之间的差距，我们引入了一种基于每流公平份额(FFS)概念的新公平模型。在每个路径节点上，流的FFS以数据包标签的形式表示，并在转发路径上更新，传达其当前出口带宽的公平份额。数据包携带的FFS与概率性转发的结合，实现了流的有效和可扩展隔离，同时具有最小开销。FFS是第一个将低实现和部署开销与有效带宽隔离相结合的系统，同时保持对源地址欺骗和DDoS攻击的鲁棒性，并提供高性能、可扩展性以及最小的延迟和抖动。我们证明FFS能够在隔离15种不同CCA的带宽的同时，保持延迟和抖动最小。我们的高速实现能够在商用硬件上维持160 Gbps的线路速率。在真实的互联网拓扑上评估时，FFS在带宽分配的中位数和总量上都优于几种最新且安全的带宽隔离系统。在我们的安全分析中，我们证明FFS为每个流量流保证了一个非零的带宽分配下限，确保即使结合源地址欺骗，DDoS攻击也无法阻止合法通信。最后，我们提出了FFS的扩展，为发送方提供准确且安全的速率反馈，允许快速速率适应且最小化数据包丢失。</span></span></p><p cid="n556" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f23-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f23-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n558" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">135、Limitless Scalability: A High-Throughput and Replica-Agnostic BFT Consensus</span></span></p><p cid="n559" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">传统的拜占庭容错（BFT）共识协议采用星型拓扑结构，由领导者处理所有消息传输，导致副本数量增加时性能迅速下降。最近，许多研究通过探索多层拓扑结构（如树结构）来减少领导的扇出，以提高可扩展性。然而，这些方法要么依赖于多项式扇出来保持容错能力，要么受到拓扑深度对吞吐量影响的限制，最终仅带来有限的可扩展性提升。为此，我们提出了Tide，这是首个能够随着副本数量增长而保持稳健性能的BFT共识协议，这得益于我们对对数扇出拓扑和高并行流水线的设计。Tide在拓扑设计中利用冗余连接作为关键洞察，在不降低弹性的情况下减少扇出。Tide进一步引入了一种新颖的流水线机制，其中层间交互动态确定提案并行度，从而将吞吐量与拓扑深度解耦。使用100台云服务器的真实实验表明，当副本数量从100扩展到1000时，最先进协议的吞吐量下降了65%至90%，延迟增加了50倍。相比之下，Tide保持了与副本数量无关的高吞吐量，约为50ktps，比其他协议高出5倍以上，而其延迟保持在0.3秒至0.4秒之间。</span></span></p><p cid="n560" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f101-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f101-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n562" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">136、LinkGuard: A Lightweight State-Aware Runtime Guard Against Link Following Attacks in Windows File System</span></span></p><p cid="n563" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Windows文件系统中的链接跟随（LF）攻击允许攻击者通过滥用精心设计的符号链接组合（链接链），将正常的文件操作悄悄重定向到受保护的文件，从而实现对受保护文件的任意操作。这类攻击通常表现为单步攻击或多步攻击，具体取决于构建的链接链的顺序。现有的针对LF攻击的防御措施要么依赖于复杂的建模，要么存在兼容性差和适用性有限的问题，且没有一种能够为不同类型的LF攻击提供全面保护。在本文中，我们提出了LinkGuard，一个针对Windows系统的轻量级状态感知运行时防御机制。LinkGuard的创新之处在于其两阶段设计：第一阶段通过执行动态主体过滤来提高防御效率，仅监控涉及链接链创建和跟随的文件操作及相关主体；第二阶段基于有限状态机（FSM）的规则匹配来精确防御LF攻击，确保有效且准确的防御。我们在五个代表性的Windows系统上评估了LinkGuard的原型，以验证其兼容性。在一个包含70个真实世界漏洞的数据集上，LinkGuard成功缓解了所有单步攻击和95.45%的多步攻击，并且在良性操作上零误报。在微基准测试中，LinkGuard平均仅产生1%的开销，在实际应用工作负载中产生3.4%的开销，同时在良性文件操作上仅增加5毫秒的延迟。</span></span></p><p cid="n564" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2943-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2943-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n566" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">137、LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline</span></span></p><p cid="n567" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">错误二分法是一项重要的安全任务，旨在理解受软件错误影响的版本范围，即确定引入错误的提交。然而，传统的基于补丁的二分方法面临几个重大障碍：例如，它们假设引入错误的提交（BIC）和补丁提交修改相同的函数，但这并非总是成立；它们通常仅依赖代码变更，而提交消息中经常包含丰富的漏洞相关信息；它们还基于简单的启发式方法（例如假设BIC初始化了补丁中删除的代码行），缺乏对漏洞的逻辑分析。在本文中，我们观察到大型语言模型（LLMs）有潜力突破现有解决方案的障碍，例如能够很好地理解补丁和提交中的文本数据和代码。我们开发了一个全面的多阶段流程，利用LLLMs来（1）充分利用完整的补丁信息，（2）让LLM评估错误的逻辑以及提交成为引入错误提交的可能性，以及（3）通过多次筛选过程逐步缩小候选范围。在我们的评估中，我们证明该方法比最先进的解决方案准确率提高38%以上。我们的结果进一步证实了全面的多阶段流程是必不可少的，因为它比简单的LLM应用准确率提高60%。</span></span></p><p cid="n568" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s990-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s990-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n570" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">138、Loki: Proactively discovering online scams by mining toxic search queries</span></span></p><p cid="n571" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在线电子商务诈骗，从购物诈骗到宠物诈骗，每年在全球造成数百万美元的经济损失。为此，安全社区已经开发了高度准确的检测系统，能够确定网站是否具有欺诈性。然而，寻找可作为输入提供给这些下游检测系统的候选诈骗网站具有挑战性：依赖用户报告本质上是被动的且反应缓慢，而主动发出搜索引擎查询以返回候选网站的系统则存在覆盖范围有限且无法推广到新型诈骗类型的问题。在本文中，我们提出了LOKI系统，该系统旨在识别可能返回大量欺诈性网站的搜索引擎查询。LOKI实现了基于特权信息学习（LUPI）和搜索引擎结果页面（SERP）特征提取的关键词评分模型。我们在10个主要诈骗类别中对LOKI进行了严格验证，并在所有类别中展示了相较于启发式和数据驱动基线20.58倍的发现率提升。利用仅包含1,663个已知诈骗网站的小型种子集，我们使用通过该方法识别的关键词发现了52,493个先前未报告的野外诈骗案例。最后，我们证明了LOKI可以推广到先前未见过的诈骗类别，突显了其在发现新兴威胁方面的实用性。</span></span></p><p cid="n572" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s184-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s184-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n574" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">139、Looma: A Low-Latency PQTLS Authentication Architecture for Cloud Applications</span></span></p><p cid="n575" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">量子计算机威胁着打破传统TLS的密码学基础，促使向后量子密码学转变。然而，后量子认证会带来显著的性能开销，特别是在高握手率的云环境中进行相互TLS认证时。我们提出了Looma，一种快速的后量子认证架构，它将认证分为快速的路径上签名/验证操作和慢速的路径外异步预计算，在不牺牲安全性的情况下减少了握手延迟。集成到TLS 1.3中，与基于Dilithium-2的基线相比，Looma将PQTLS握手延迟降低了最多44%。我们的研究结果表明，Looma在云环境中扩展后量子安全通信具有实用性。</span></span></p><p cid="n576" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f74-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f74-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n578" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">140、Losing the Beat: Understanding and Mitigating Desynchronization Risks in Container Isolation</span></span></p><p cid="n579" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">当今容器提供的隔离是通过高度协调地利用Linux命名空间和cgroups实现的。然而，随着计算范式的演变，特别是对跨命名空间资源共享有强烈需求的无服务器计算的出现，这种容器保护的基础已经动摇。这种共享削弱了容器的隔离模型，正如我们在研究中发现的，导致了命名空间-cgroup不同步（NCD）漏洞的出现。在本文中，我们对此类风险进行了研究，旨在确定其根本原因并理解其影响。我们的研究揭示，流行的容器工具都存在NCD风险，这在我们发现的四个新漏洞和一个错误中得到了证实。从根本上说，命名空间共享扩展了容器的隔离边界，这可能违反cgroups设定的限制，从而削弱了这两种机制提供的联合保护。这种冲突通常无法通过现有的容器工具调和。为了应对这一挑战并满足命名空间共享的需求，我们提出了一个内核级解决方案，以统一命名空间和cgroups在监控容器实例资源方面的分散职责。我们的设计将命名空间处理的资源管理与cgroups强制执行的限制相结合，并确定了它们应遵循的协作策略。分析和评估表明，我们的方法有效缓解了NCD风险，同时对Linux内核、主流容器工具和实际应用程序造成的成本可以忽略不计，并保持与这些系统的完全兼容性。</span></span></p><p cid="n580" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1381-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1381-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n582" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">141、Mapping the Cloud: A Mixed-Methods Study of Cloud Security and Privacy Configuration Challenges</span></span></p><p cid="n583" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">云服务配置错误仍然是安全和隐私事件的主要原因，这通常源于云平台配置的复杂性。为了更好地理解这些挑战，我们分析了从2008年到2024年间约251,900条与安全和隐私相关的Stack Overflow帖子。通过使用主题建模和定性分析，我们系统地映射了云用例与其相关的安全和隐私配置挑战，揭示了云运营商所面临障碍的全景图。我们确定了技术性和以人为中心的问题，包括与文档不足以及缺乏针对运营商环境的上下文感知工具相关的问题。值得注意的是，身份验证和访问控制挑战出现在所有已识别的用例中，贯穿云部署、集成和维护的几乎所有阶段。我们的研究结果强调了需要可用、定制化和上下文敏感的支持工具和资源，以帮助开发人员安全地配置云服务。</span></span></p><p cid="n584" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1302-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1302-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n586" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">142、Memory Backdoor Attacks on Neural Networks</span></span></p><p cid="n587" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Torsten Krauß（维尔茨堡大学），Alexandra Dmitrienko（维尔茨堡大学），Yisroel Mirsky（内盖夫本古里安大学）</span></span></p><p cid="n588" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">神经网络通常在专有数据集上进行训练，使其成为有吸引力的攻击目标。我们提出了一种新颖的数据集提取方法，利用创新的训练时后门攻击，使恶意联邦学习（FL）服务器能够通过简单的索引过程系统性地、确定性地提取完整的客户端训练样本。与先前技术不同，我们的方法确保精确的数据恢复，而非概率性重建或幻觉，提供对记忆哪些样本及数量的精确控制，并展现出高容量和鲁棒性。受感染模型在接收到基于模式的索引触发器时会输出数据样本，从而在不影响全局模型效用的情况下，系统性地从每个客户端的本地数据中提取有意义的片段。为解决模型输出尺寸较小的问题，我们提取片段后将其重新组合。该攻击仅需对训练代码进行微小修改，可在客户端验证过程中轻易逃避检测。因此，这种漏洞代表了FL供应链的真实威胁，恶意服务器可向客户端分发修改后的训练代码，并从其更新中恢复私人数据。在分类器、分割模型和大型语言模型上的评估表明，可以从客户端模型中恢复数千个敏感训练样本，同时对任务性能影响最小，经过多轮FL后可窃取客户端的整个数据集。例如，医疗分割数据集的提取仅需3%的效用下降。这些研究结果揭示了FL系统中的关键隐私漏洞，强调了分布式训练管道中需要更强的完整性和透明度。</span></span></p><p cid="n589" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1870-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1870-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n591" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">143、Memory Band-Aid: A Principled Rowhammer Defense-in-Depth</span></span></p><p cid="n592" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DRAM中的Rowhammer位翻转使软件攻击者能够完全攻破各类系统。硬件缓解措施可以精确且高效，但它们面临漫长的部署周期和非常有限或无更新能力的缺点。因此，改进的攻击方法已多次绕过已部署的硬件保护措施，使得商用系统容易受到Rowhammer攻击。在本文中，我们提出了Memory Band-Aid，一种针对Rowhammer的纵深防御方案。Memory Band-Aid并非长期高效的硬件缓解措施的替代品，而是一种纵深防御，在硬件缓解措施对特定系统世代不足时激活。为此，Memory Band-Aid在内存控制器中引入了按线程和按存储库的DRAM访问速率限制，确保无法达到Rowhammer位翻转所需的最小行激活次数。我们在Ubuntu Linux上实现了Memory Band-Aid的概念验证，并在2个Intel和2个AMD系统上进行了测试，由于当前硬件缺乏按存储库的限制，我们基于全局带宽限制进行实现。使用这个PoC，我们发现包含少量硬件更改的完整实现在一组真实的Phoronix宏基准测试中开销仅为0%至9.4%。在导致DRAM压力的微基准测试中，我们观察到1至5.1倍的减速。这两种开销仅适用于不可信的、被限制的工作负载，例如所有用户空间程序或仅选定的沙箱，如浏览器中的沙箱。特别是因为Memory Band-Aid可以按需启用，我们得出结论，Memory Band-Aid是一种重要的纵深防御，应作为第二防御层在实际中部署。</span></span></p><p cid="n593" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s156-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s156-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n595" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">144、MEVisor: High-Throughput MEV Discovery in DEXs with GPU Parallelism</span></span></p><p cid="n596" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">去中心化金融（DeFi）是区块链上新兴的金融服务，能够实现自动和匿名的交易。在DeFi中，去中心化交易所（DEXs）维护一对代币的储备，并确定代币交换的汇率。然而，DEXs也为最大可提取价值（MEV）创造了机会，攻击者可以通过包含、排除或重新排序DEX交易来利用代币价格差异并获取利润。发现MEV机会需要高吞吐量，因为12秒的区块间隔和庞大的搜索空间施加了严格的时间限制。然而，现有工具由于依赖CPU绑定执行，频繁的状态分叉和缓慢的DEX执行，导致吞吐量低下。在本文中，我们首次利用GPU并行计算能力来提升套利和三明治策略中的MEV搜索吞吐量。更准确地说，我们将MEV机器人编译为GPU应用程序，然后启动数千个GPU线程并行搜索利润。为此，我们设计了新的解决方案来解决三大挑战：设计在GPU上模拟交易的作弊代码，提出减少GPU内存使用的内存管理器，以及设计策略感知的变异以提高输入多样性。我们实现了一个名为MeVisor的原型，它在GPU上运行DEXs，并使用并行遗传算法搜索MEV。基于以太坊的3,941个真实MEV案例进行评估，MeVisor实现了每秒330万至510万笔交易的吞吐量，比CPU基准性能高出10万倍。在2025年第一季度的大规模研究中，MeVisor估计MEV机会在2到14笔交易之间，最多可获得110万美元的MEV利润。</span></span></p><p cid="n597" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f93-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f93-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n599" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">145、MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness</span></span></p><p cid="n600" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">视觉变换器（ViTs）已成为基础架构，并作为现代视觉-语言模型的骨干网络。尽管它们表现出色，但ViTs对逃避攻击表现出明显的脆弱性，这需要开发专门针对其独特架构的对抗训练（AT）策略。虽然直接解决方案可能涉及将现有的AT方法应用于ViTs，但我们的分析揭示了显著的不兼容性，特别是与最先进（SOTA）方法（如Generalist（CVPR 2023）和DBAT（USENIX Security 2024））存在明显差异。本文对ViTs中的对抗鲁棒性进行了系统研究，并基于其自编码器自监督预训练提供了新的互信息（MI）理论分析。具体而言，我们证明了在基于ViT的自编码器中，对抗样本与其潜在表示之间的MI应通过推导出的MI边界进行约束。基于这一见解，我们提出了一个名为MIMIR的自监督AT方法，该方法采用MI惩罚机制，通过自编码器的掩码图像建模促进对抗预训练。在CIFAR-10、Tiny-ImageNet和ImageNet-1K上的大量实验表明，MIMIR能够持续提高自然准确率和鲁棒性，在ImageNet-1K上超越了最先进的AT结果。值得注意的是，MIMIR对未预见攻击和常见损坏数据表现出更强的鲁棒性，并且能够抵御拥有完整防御机制知识的自适应攻击。我们的代码和训练模型可在以下公开获取：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/xiaoyunxxy/MIMIR" target="_blank">https://github.com/xiaoyunxxy/MIMIR</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">。</span></span></p><p cid="n601" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1813-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1813-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n603" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">146、MinBucket MPSI: Breaking the Max-Size Bottleneck in Multi-Party Private Set Intersection</span></span></p><p cid="n604" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">多方隐私集合求交（基数）协议使T（T&gt;2）方，每方持有一个私有集合，能够联合计算集合的交集（或基数），而不会向其他方泄露任何额外信息。迄今为止，所有已知的MPSI（MPSI-Card）协议都需要与大规模集合大小成比例的通信复杂度，这从根本上阻碍了它们在具有异构输入规模的实际应用中的高效部署。在这项工作中，我们提出了一种基于新协议的MPSI新框架：批量成员条件随机生成和联合私有相等性测试。通过实例化这一框架，我们开发了两种MPSI协议，其通信复杂度与小集合的大小成线性关系，与大集合的大小成对数关系。一种协议可抵御任意数量的合谋方，而另一种协议可抵御(T-2)个合谋方。此外，我们还开发了一种称为联合置换私有相等性测试的协议，并提出了MPSI-Card框架。通过实例化这一框架，我们推导出一种具有类似通信效率的MPSI-Card协议：与小集合大小成线性关系，与大集合大小成对数关系，可抵御任意数量的合谋方。我们在局域网和广域网环境中实现了我们的协议并进行了广泛实验。实验结果表明，随着集合间大小差异或持有小集合的参与者数量的增加，我们的协议实现了显著更好的性能。在5个持有大规模集合（大小为2^20）和5个持有小规模集合（大小为2^10）的参与方设置中，使用单线程和10 Mbps带宽，我们的MPSI（MPSI-Card）协议仅需12.2（12.2）MB的通信量和129.86（130.05）秒的运行时间。与Wao等人（USENIX Security 2024）的最先进MPSI和高等人（PETS 2024）的MPSI-Card相比，我们的协议实现了通信成本降低157倍（76倍）和运行时间加速12.7倍（3.1倍）。</span></span></p><p cid="n605" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f182-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f182-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n607" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">147、Mirage: Private, Mobility-based Routing for Censorship Evasion</span></span></p><p cid="n609" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在专制和高度监控的环境中，传统通信网络容易受到审查、监控和破坏。虽然像Tor这样的去中心化匿名网络能提供强大的隐私保障，但它们仍然依赖于集中式互联网基础设施，使其容易受到大规模封锁或关闭。为解决这些局限性，我们提出了MIRAGE，一种基于移动性的隐私保护消息系统，专为抗审查通信而设计。MIRAGE采用基于区域的路由方案，根据人群的高层移动模式概率性地转发消息。为防止个人移动行为的泄露，MIRAGE通过局部差分隐私保护用户的移动模式，确保参与网络不会通过可观察的路由决策揭示个人的位置历史。我们在Cadence中实现了MIRAGE，这是一个开源模拟器，提供了一个统一框架，用于使用节点间随时间推移的近似地理遭遇来评估基于移动性的协议。我们分析了MIRAGE的隐私与效率权衡，并使用真实世界轨迹对其性能进行了评估：(1)传统流行病和基于随机游走的路由协议，以及(2)最先进的隐私保护地理路由协议。这些轨迹包括：一个是在不同城市地点收集的行人移动模式，另一个是出租车运营的GPS轨迹。我们的结果表明，与流行病路由相比，MIRAGE显著减少了消息开销，在投递率方面优于概率泛洪，同时比现有技术提供更强的隐私保障。</span></span></p><p cid="n610" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s237-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s237-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n612" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">148、Mobius: Enabling Byzantine-Resilient Single Secret Leader Election with Uniquely Verifiable State</span></span></p><p cid="n613" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">单密钥领导者选举（SSLE）协议能够在确保不可预测性的情况下，在一组注册节点中每轮选举出一个单一领导者。以太坊已将SSLE确定为其发展路线图中的关键组成部分，并采用其作为应对潜在攻击的潜在解决方案。然而，我们识别出一种新型攻击，称为&#34;状态唯一性&#34;攻击，该攻击由恶意领导者提出多个可公开验证的状态引起。这种攻击破坏了后续领导者选举中的&#34;唯一性&#34;属性，并很可能导致上层协议基本安全属性（如活跃性）的违反。这一漏洞源于将唯一性保证降低为每次选举只有一个状态的设计，并可推广到现有的SSLE构造中。我们基于理论分析和在以太坊上的实际执行进一步量化了这种攻击的严重性，强调了设计可证明安全的SSLE协议所面临的严峻挑战。为了解决&#34;状态唯一性&#34;攻击同时确保安全性和实际性能，我们提出了一个名为Mobius的通用SSLE协议，该协议不依赖额外的信任假设。具体而言，Mobius防止每次选举生成多个可验证状态，并通过创新的&#34;近似唯一随机化&#34;机制在连续执行中实现唯一状态。除了在通用可组合性框架中提供全面的安全分析外，我们还开发了Mobius的概念验证实现，并进行了广泛实验以评估其安全性和开销。实验结果表明，Mobius在显著降低协议执行过程中的通信复杂度的同时增强了安全性，在注册阶段实现了超过80%的减少。</span></span></p><p cid="n614" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2407-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2407-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n616" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">149、MUTATO: Enhancing Fuzz Drivers with Adaptive API Option Mutation</span></span></p><p cid="n617" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">模糊测试是发现漏洞和提高软件系统可靠性的核心技术。最近的研究表明，现代覆盖率引导模糊测试的主要瓶颈不在于模糊测试工具本身，而在于模糊驱动程序的构建——特别是它们在探索库API中选项参数时的有限灵活性。现有方法主要关注变异输入数据，常常忽略了从根本上影响API行为并可能隐藏关键漏洞的配置选项。为解决这一差距，我们提出了MUTATO，一种新的多维度模糊驱动程序增强方法，它使用覆盖率引导的ε-贪婪策略系统且自适应地变异输入数据和选项参数。与需要侵入性修改模糊测试工具或仅针对程序级选项的先前工作不同，MUTATO在驱动程序级别运行，确保了与模糊测试工具无关的适用性，并能与手动和自动生成的驱动程序无缝集成。我们进一步引入了选项参数模糊测试语言（OPFL）来指导驱动程序的增强。在10个广泛使用的C/C++库上进行的大量实验表明，与原始AFL++和LibFuzzer驱动程序相比，MUTATO增强的驱动程序平均分别实现了14%和13%的代码覆盖率提升，并发现了12个先前未知的漏洞，其中包括3个CVE。值得注意的是，我们在API中发现了4个漏洞，而OSS-Fuzz尽管进行了超过18,060小时的模糊测试却未能检测到这些漏洞。</span></span></p><p cid="n618" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s820-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s820-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n620" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">150、MVP-ORAM: a Wait-free Concurrent ORAM for Confidential BFT Storage</span></span></p><p cid="n621" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">众所周知，仅靠加密不足以保护数据隐私。执行操作时暴露的访问模式也可能被用于推理攻击。 oblivious RAM (ORAM) 通过使客户端请求变得不可知来隐藏访问模式。然而，现有协议在支持并发客户端和拜占庭容错(BFT)方面仍然存在局限性。我们提出了 MVP-ORAM，这是第一个支持易故障并发客户端的无等待 ORAM 协议。与之前的工作不同，MVP-ORAM 避免使用需要额外安全假设的可信代理，以及基于客户端间通信或分布式锁的并发控制机制，这些机制限制了整体吞吐量和容忍故障客户端的能力。相反，MVP-ORAM 使客户端能够执行并发请求并在发生时合并冲突更新，满足无等待特性，即客户端独立于其他客户端的性能或故障取得进展。由于等待和冲突自由是根本上矛盾的目标，无法在异步并发 ORAM 服务中同时实现，我们定义了一个依赖于应用程序工作负载和并发客户端数量的较弱不可知性概念，并证明 MVP-ORAM 在客户端执行倾斜块访问的实际场景中是安全的。通过实现无等待特性，MVP-ORAM 可以无缝集成到现有的机密 BFT 数据存储中，创建了第一个 BFT ORAM 构造。我们在一个机密 BFT 数据存储上实现了 MVP-ORAM，并证明我们的原型在现代云环境中每秒可以处理数百次 4KB 访问。</span></span></p><p cid="n622" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1809-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1809-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n624" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">151、MVPNalyzer: An Investigative Framework for Auditing the Security </span></span><span md-inline="html_entity" data-content="&amp;" style="box-sizing: border-box;"><span leaf="">&amp;</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf=""> Privacy of Mobile VPNs</span></span></p><p cid="n625" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">移动用户越来越依赖虚拟专用网络(VPN)来保护自己免受跟踪、监控和审查。VPN应用通过要求拦截用户流量而处于特权地位。虽然这保护了终端用户流量免受恶意网络中介(如监控互联网服务提供商)的侵害，但它导致了一种关键的&#34;信任转移&#34;，即从这些网络中介转移到VPN提供商。然而，尽管这一角色至关重要，但VPN应用，尤其是在移动平台上，仍然缺乏充分的审计。在这项工作中，我们提出了MVPN-Audit，一个可扩展的框架，用于系统分析Android VPN应用。该框架旨在处理Android VPN生态系统的独特挑战，使能够对VPN应用在网络各层的行为进行详细调查。我们将我们的框架应用于Google Play商店中的281个流行VPN应用，并发现了基本和关键问题：61个应用传输未加密数据，其中5个以明文形式发送敏感的VPN配置文件，允许攻击者劫持VPN隧道连接；29个应用将用户流量(包括DNS)泄露到隧道外；169个应用未能混淆流量以避免简单阻塞；76个应用传输广告ID，这是一种广泛用于设备和用户跟踪的设备唯一标识符；107个应用在其VPN配置文件中未能实施最佳安全实践。这些应用的总安装量达数亿次，突显了受影响用户的规模。我们的研究结果揭示了开发者疏忽的令人担忧的模式，突显了执行不力、透明度不足和维护不善如何继续削弱基本的安全保障。</span></span></p><p cid="n626" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1573-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1573-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n628" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">152、NetCap: Data-Plane Capability-Based Defense Against Token Theft in Network Access</span></span></p><p cid="n629" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">令牌在企业网络访问控制中发挥着至关重要的作用，它通过在各种协议（如JSON Web令牌、OAuth 2.0）中实现安全的身份验证和授权，使用户能够使用有效的访问令牌访问授权资源，而无需重复提交凭据。然而，授权主机内所有进程所获得的普遍信任，加上令牌的较长生命周期，为恶意进程劫持令牌并冒充合法用户创造了机会。这种威胁影响广泛范围的协议，并已导致众多真实世界事件。在本文中，我们提出了NetCap，这是一种新的防御机制，旨在防止攻击者在企业环境中使用被盗令牌访问未授权资源。其核心思想是引入不可伪造的、进程级的能力，这些能力与授权进程绑定。这些能力被持续嵌入到进程的网络流量中，以供目标资源进行验证，并且频繁刷新。进程身份与能力之间的这种绑定确保了即使访问令牌被恶意进程窃取，没有有效能力也无法通过身份验证。为了支持网络中进程生成的大量请求，NetCap引入了一种基于可编程交换机和eBPF的新型数据平面设计。通过多种优化技术，我们的系统支持能力的内联生成和嵌入，使大量流量能够以线路速率处理且开销极小。我们的广泛评估表明，NetCap在各种协议和实际应用中保持线路速率的网络性能，同时开销可忽略不计，并有效保护这些应用程序免受令牌盗窃攻击。</span></span></p><p cid="n630" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f273-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f273-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n632" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">153、NetRadar: Enabling Robust Carpet Bombing DDoS Detection</span></span></p><p cid="n633" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">地毯式轰炸攻击是分布式拒绝服务（DDoS）攻击日益普遍的一种变体，它同时攻击受害者网络中的多个服务器，通过最小化每流恶意流量吞吐量来规避检测。聚合的恶意流量压垮了网络接入点（如网关），导致拒绝服务。此外，高级攻击者采用应用层攻击方法生成在语义和流量体积上都不明显的恶意流量，使得现有的DDoS检测机制失效。我们提出了NetRadar，一个能够实现准确且稳健的地毯式轰炸攻击检测的DDoS检测器。NetRadar利用服务器-网关协作架构，聚合从受害者网络收集的流量和服务器端特征，并进行跨服务器分析以定位受害服务器。为实现服务器辅助的地毯式轰炸检测，我们引入了一个兼容多种服务的通用服务器端特征集，以及一种能够处理运行时特征不匹配问题的稳健模型训练方法。此外，我们还提出了一种高效的跨服务器入站流量分析方法，有效利用了地毯式轰炸流量的相似性，同时降低了计算开销。在真实和模拟数据集上的评估表明，NetRadar的检测性能优于最先进的解决方案，在所有地毯式轰炸检测场景中均实现了超过94%的准确率。</span></span></p><p cid="n634" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2118-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2118-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n636" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">154、NeuroStrike: Neuron-Level Attacks on Aligned LLMs</span></span></p><p cid="n637" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">安全对齐对于大型语言模型（LLM）的道德部署至关重要，它引导模型避免生成有害或不道德的内容。当前的对齐技术，如监督微调和基于人类反馈的强化学习，仍然存在脆弱性，可以通过精心设计的对抗性提示绕过。不幸的是，此类攻击依赖于反复试验，缺乏跨模型的泛化能力，且受可扩展性和可靠性的限制。本文提出了NeuroStrike，一种新颖且可泛化的攻击框架，它利用了对齐技术引入的根本性漏洞：对稀疏、专门的安全神经元的依赖，这些神经元负责检测和抑制有害输入。我们将NeuroStrike应用于白盒和黑盒场景：在白盒场景中，NeuroStrike通过前馈激活分析识别安全神经元，并在推理过程中剪除它们以禁用安全机制。在黑盒场景中，我们提出了首个LLM画像攻击，利用安全神经元的可转移性，在开源权重代理模型上训练对抗性提示生成器，然后将其部署到黑盒和专有目标模型上。我们在来自主要LLM开发者的20多个开源权重LLM上评估了NeuroStrike。通过移除目标层中不到0.6%的神经元，NeuroStrike仅使用普通恶意提示就实现了76.9%的平均攻击成功率（ASR）。此外，NeuroStrike泛化到四个多模态LLM，对不安全图像输入的攻击成功率达到100%。安全神经元在不同架构间有效转移，使11个微调模型和5个蒸馏模型的ASR分别提高到78.5%和77.7%。黑盒LLM画像攻击在五个黑盒模型（包括谷歌的Gemini系列）上实现了63.7%的平均ASR。</span></span></p><p cid="n638" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s660-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s660-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n640" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">155、NEXUS: Towards Accurate and Scalable Mapping between Vulnerabilities and Attack Techniques</span></span></p><p cid="n641" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">通用漏洞与暴露(CVE)计划每年记录数千个已知漏洞，但没有关于这些漏洞可能如何被攻击者利用的可操作上下文。另一方面，MITRE ATT&amp;CK框架概述了攻击战术、技术和程序(TTPs)，但没有将其与特定漏洞联系起来。虽然实现CVE描述到TTPs的自动映射可以允许更准确、更高效地检测和缓解威胁，但现有工作面临几个挑战：(i)缺乏将CVEs与TTPs链接的大规模、高质量数据集；(ii)现有数据中存在数据分布不均和关键TTPs缺失的问题；(iii)从非结构化CVE描述中准确提取敌对行为的困难；以及(iv)缺乏用于持续修正映射的自适应学习机制。本文通过NEXUS框架解决了这些挑战，该框架可自动将CVEs映射到TTPs。我们的评估(基于一个新构建的数据集，涵盖208个TTPs和92K+个CVEs，以及其他公共数据集)表明，NEXUS在CVE到TTP的映射中实现了97.94%的最高F1分数，并且能够处理新的CVE条目，而现有工作的最高F1分数仅为67.68%。</span></span></p><p cid="n642" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2926-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2926-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n644" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">156、ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data</span></span></p><p cid="n645" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">提示注入攻击旨在污染大语言模型(LLM)的输入数据，误导其完成攻击者选择的任务而非预期任务。在许多应用和代理中，输入数据来源于多个来源，每个来源贡献整体输入的一部分。在这些多源场景中，攻击者可能只控制部分来源并污染相应片段，但通常不知道这些片段在输入中的排列顺序。现有的提示注入攻击要么假设整个输入数据来自攻击者控制的单一来源，要么忽略不同来源片段排列的不确定性。因此，它们在涉及多源数据的领域中效果有限。在这项工作中，我们提出了ObliInjection，这是首个针对具有多源输入数据的大语言模型应用和代理的提示注入攻击。ObliInjection引入了两项关键技术创新：顺序无关损失(order-oblivious loss)，用于量化无论干净和污染片段如何排列，大语言模型都会完成攻击者选择任务的可能性；以及顺序GCG算法(orderGCG algorithm)，专门用于最小化顺序无关损失并优化污染片段。跨越三个不同应用领域数据集和十二个大语言模型的全面实验表明，ObliInjection高度有效，即使输入数据中只有6-100个片段中的一个被污染。我们的代码和数据可在以下网址获取：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/ReachalWang/ObliInjection" target="_blank">https://github.com/ReachalWang/ObliInjection</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">。</span></span></p><p cid="n646" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f702-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f702-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n648" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">157、OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity</span></span></p><p cid="n649" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">近年来，分支目标缓冲器（BTB）在系统安全研究中引起了广泛关注。在某些攻击场景中，该组件在逻辑或物理上被共享，被攻击者滥用以构建侧信道，从而泄露受害进程的敏感分支信息。然而，现有的BTB侧信道攻击要么因跨权限隔离机制而无法从用户模式泄露内核控制流信息，要么在分支监控中存在空间分辨率有限的问题。在本文中，我们提出了Occupy+Probe，一种基于驱逐的新型BTB侧信道攻击，它通过直接从用户模式成功暴露内核控制流行为来弥合这些差距。我们的方法从对Intel处理器上与偏移量相关的BTB更新机制的深入逆向工程开始，并揭示&#34;在用户模式下创建的BTB条目可以直接被内核模式条目替换，而不管底层的替换策略和硬件隔离如何&#34;，这构成了Occupy+Probe的基础。与现有的BTB侧信道攻击相比，Occupy+Probe消除了攻击者和受害者之间条目共享的需求。此外，它在分支监控中实现了指令级别的粒度，超越了现有基于驱逐的BTB侧信道器的空间分辨率。我们通过实验证明，Occupy+Probe可以在各种Intel处理器上以高空间分辨率跨权限边界泄露控制流信息。此外，我们通过针对Linux内核加密API的详细案例研究验证了Occupy+Probe的实际有效性，展示了其破坏关键内核操作的潜力。此外，与先前基于驱逐的BTB侧信道相比，Occupy+Probe展示了提取内核分支标签值的独特能力，这可用于破坏KASLR。</span></span></p><p cid="n650" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s925-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s925-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n652" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">158、Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography</span></span></p><p cid="n653" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">多模态大语言模型（MLLMs）将文本与其他模态（如图像）相结合，展现出强大的能力，并在现实商业系统中日益广泛应用。然而，其日益普及也引发了关于滥用的担忧，例如生成有害内容。为缓解这些风险，对齐技术常被用于使模型行为与人类价值观保持一致。尽管有这些努力，但近期研究表明，越狱攻击可以绕过对齐并引发不安全输出。目前，大多数现有越狱方法针对开源模型设计，对集成额外过滤器的商业MLLM系统的效果有限。这些过滤器能够检测并阻止恶意输入和输出内容，显著降低越狱威胁。本文揭示，这些安全过滤器的成功严重依赖于一个关键假设：恶意内容必须在输入或输出中明确可见。这一假设在传统LLM集成系统中通常有效，但在MLLM集成系统中却不再成立，因为攻击者可以利用多种模态来隐藏对抗意图，导致现有MLLM集成系统产生虚假的安全感。为挑战这一假设，我们提出了Odysseus，一种新颖的越狱范式，引入双重隐写术，将恶意查询和响应 covertly 嵌入到看似无害的图像中。我们的方法通过四个阶段进行：（1）恶意查询编码，（2）隐写术嵌入，（3）模型交互，和（4）响应提取。我们首先将攻击者指定的恶意提示编码为二进制矩阵，并使用隐写术模型将其嵌入图像中。修改后的图像将被输入到目标MLLM集成系统中。我们鼓励目标MLLM集成系统将生成的不当内容植入到载体图像中（通过隐写术），供攻击者本地解码隐藏的响应。在基准数据集上的广泛实验表明，我们的Odysseus成功攻击了多个前沿且现实的MLLM集成系统，包括GPT-4o、Gemini-2.0-pro、Gemini-2.0-flash和Grok-3，攻击成功率高达99%。它暴露了现有防御中的一个根本盲点，呼吁重新思考MLLM集成系统中的跨模态安全问题。</span></span></p><p cid="n654" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f808-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f808-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n656" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">159、On Borrowed Time: Measurement-Informed Understanding of the NTP Pool’s Robustness to Monopoly Attacks</span></span></p><p cid="n657" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">互联网服务和应用严重依赖于网络时间的可用性和准确性。网络时间协议（NTP）是最古老的核心网络协议之一，至今仍是互联网上时钟同步的实际标准机制。尽管存在多个NTP基础设施，但其中之一&#34;NTP池&#34;因其两个基本原因而成为一个极具吸引力的攻击目标：1）它采用分布式管理，基于志愿者服务器；2）被广泛使用，包括全球的物联网和基础设施设备。我们首次收集了关于NTP池的直接、非推断性和全面的数据，包括：纵向的服务器和账户成员资格、服务器配置、时间质量、别名和全球查询流量负载。我们在九个月内收集了完整且细粒度的数据，发现了超过15,000台服务器（包括活跃和非活跃服务器），并对NTP池的使用情况、动态性和稳健性提供了新的见解。通过分析地址别名、账户和网络连接，我们发现池中只有19.7%的活跃服务器是完全独立的。最后，我们证明，拥有我们数据的攻击者能够更好、更精确地发起&#34;垄断攻击&#34;，只需10台或更少的恶意NTP服务器，就能捕获90%国家中绝大多数NTP池流量。我们的研究结果提出了多种可以改进该池稳健性的途径。</span></span></p><p cid="n658" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f541-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f541-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n660" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">160、On the Security Risks of Memory Adaptation and Augmentation in Data-plane DoS Mitigation</span></span></p><p cid="n662" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">商用交换机的数据平面可编程性正在通过支持自适应的线速率缓解策略，重塑拒绝服务（DoS）防御的格局。最近的系统如Cerberus [SP&#39;24]利用控制平面支持来扩展有限的交换机内存，从而能够快速应对不断演变的攻击。在本文中，我们揭示了该模型中一个微妙但关键的漏洞；即，正是那些使防御系统具有敏捷性和可扩展性的机制，可能被一类新的协调式DoS攻击所破坏。我们提出了Heracles，这是首个利用可编程交换机中的硬件级约束来协调数据平面和控制平面内存中精确资源争用的攻击。通过利用侧信道时序信号，Heracles触发了同步增强、内存挤压和时间窗口利用，这是三种正交的争用策略，会显著降低甚至完全禁用DoS缓解能力。我们在真实的Tofino硬件上实现并测试了Heracles，表明它可以可靠地破坏各种DoS攻击配置文件下的DoS防御，即使使用松散（1-2秒）时间同步的攻击源也是如此。为缓解这一威胁，我们提出了Shield，一种多层DoS缓解草图架构，它解耦了控制平面和数据平面层的内存操作，在保持线速率性能和检测准确性的同时，有效缓解了Heracles攻击。</span></span></p><p cid="n663" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1857-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1857-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n665" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">161、One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases</span></span></p><p cid="n666" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">电子邮件地址作为在线账户管理的通用标识符，其别名机制在电子邮件提供商与外部平台之间引入了显著的身份混淆问题。本文首次对电子邮件别名引起的不一致性问题进行了系统性分析，其中提供商将别名地址（如</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf="">ALICE@example.com</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">、</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf="">alice+work@example.com</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">）视为基础电子邮件（</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf="">alice@example.com</span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">）的额外入口，而平台通常将它们视为不同的身份。通过对28家电子邮件提供商和18个在线平台的别名机制进行实证评估，我们揭示了关键差距：（1）只有Gmail完整记录了其别名规则，而11家提供商默默地支持未文档化的别名行为；（2）由于缺乏标准化文档和实际实施，平台要么无法区分别名地址，要么过度激进地排除了包含特定符号的所有电子邮件。真实世界的滥用案例表明，攻击者利用别名在npm中从单个基础电子邮件创建多达139个账户用于垃圾邮件活动。我们的用户研究进一步强调了安全风险，显示31.65%具有别名知识的参与者因提供商实现不一致而将钓鱼邮件误认为是合法的电子邮件别名。那些认为自己理解电子邮件别名的用户，特别是受教育程度高、男性和技术参与者，更容易受到钓鱼攻击。我们的研究结果强调了电子邮件别名标准化和透明的迫切需求。我们贡献了OriginMail工具，帮助平台解决别名混淆问题，并向相关利益相关者披露漏洞。</span></span></p><p cid="n667" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s148-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s148-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n669" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">162、OptiMix: Scalable and Distributed Approaches for Latency Optimization in Modern Mixnets</span></span></p><p cid="n670" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">提出的混网（mixnets）提供网络级匿名性，但代价是增加了通信延迟，这 consequently限制了它们仅适用于延迟容忍型应用，缩小了参与此类用例的客户端匿名集合。解决这一问题需要优化延迟，正如最近在LARMix（NDSS&#39;24）和LAMP（NDSS&#39;25）中通过节点排列和战略路由所探索的那样。然而，这些方法针对特定的混网设计，依赖简化的模型和信任假设，或存在实际效率有限的问题。相比之下，OptiMix通过引入一种通用的低延迟混网模型来弥合这些差距，该模型可适应所有成熟的设计。为此，我们首先提出了一种高效的分布式协议，用于在混网中排列节点，在保持对对手的无偏性（unbiasability）的同时实现低延迟特性。其次，我们引入了优化通信延迟的新型战略路由方案。第三，我们设计了一种负载均衡算法，能够均匀分配流量而不损害路由策略的延迟优化特性。第四，我们使用已部署的Nym混网数据进行了广泛评估，展示了在各种混网设计中显著降低延迟的同时最小化匿名损失——与最先进的解决方案相比实现了高达4倍的性能提升。最后，考虑到延迟减少会导致匿名性降低或带宽开销增加——正如匿名三难困境（anonymity trilemma）所述——我们提出了一种覆盖路由机制，使客户端能够受益于低延迟混网而不损害匿名性，代价是生成额外的覆盖流量。</span></span></p><p cid="n671" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s2680-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s2680-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n673" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">163、OSAVRoute: Advancing Outbound Source Address Validation Deployment Detection with Non-Cooperative Measurement</span></span></p><p cid="n674" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">源IP地址欺骗促进了各种恶意攻击，而出站源地址验证（OSAV）仍然是防止欺骗数据包离开网络的最佳当前实践。准确测量OSAV的部署对于研究网络对IP欺骗的脆弱性至关重要。然而，此类测量通常需要从被测试网络内部发送欺骗数据包，需要网络运营商的合作。本文介绍了OSAVRoute，这是首个能够捕获OSAV部署细粒度特征的非合作系统。与现有只能识别OSAV缺失的非合作方法不同，OSAVRoute能够识别OSAV的存在与缺失，并进一步测量其阻塞粒度和阻塞深度，实现了先前仅限于合作方法的能力。OSAVRoute通过显式追踪欺骗数据包的转发路径实现这一功能，能够识别其生成和传播行为。OSAVRoute的准确率达到99.4%，覆盖范围比CAIDA Spoofer多3.1倍的自治系统（AS），它揭示84.2%的测试AS未部署OSAV，尤其是在ISP网络中。在实施OSAV的网络中，95.5%在前两个IP跳内阻塞欺骗数据包，但表现出各种阻塞粒度，其中/22到/24最为常见。此外，我们首次揭示了MANRS参与与OSAV部署之间的正相关关系。</span></span></p><p cid="n675" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s17-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s17-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n677" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">164、PACS: Privacy-Preserving Attribute-Driven Community Search over Attributed Graphs</span></span></p><p cid="n678" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在数据驱动应用中，属性驱动的社区搜索已引起越来越多的关注，旨在帮助用户在属性图上找到满足特定要求的高质量子图。然而，在进行社区搜索时，很少有工作考虑数据隐私。一个关键原因是现实世界中的图规模持续增长，而属性驱动的社区搜索涉及在加密图数据上计算复杂指标，包括结构凝聚性和属性相关性，这些计算过于耗时，难以实际应用。本文首次提出了一种面向云的隐私保护属性驱动社区搜索实用方案，命名为PACS。PACS使服务器能够在接近毫秒的时间内高效响应属性驱动的社区搜索，同时无需访问属性图和搜索结果的敏感信息。为此，我们设计了两种结构：安全社区索引和安全边表，用于保护原始属性图的隐私。安全社区索引使云服务器能够高效识别满足结构凝聚性且具有最高属性分数的目标社区。特别是，我们采用内积加密来基于加密属性向量评估社区的属性驱动分数。通过BGN同态加密构建的安全边表，使云服务器能够安全地检索目标社区的边信息而无需了解其细节。我们进行了全面的安全分析，证明PACS实现了CQA2安全性。在真实社交网络数据集上的实验评估表明，PACS在处理属性驱动社区搜索时实现了接近毫秒级的效率。</span></span></p><p cid="n679" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1586-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1586-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n681" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">165、Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm</span></span></p><p cid="n682" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着大型语言模型的快速发展，其恶意使用（特别是在生成钓鱼内容方面）的潜在威胁日益普遍。利用LLM的能力，恶意用户可以合成没有拼写错误和其他易于检测特征的钓鱼邮件。此外，这类模型能够生成针对特定主题的钓鱼信息，根据目标领域定制内容，提高成功率。由于LLM生成的钓鱼邮件通常缺乏清晰或可区分的语言特征，检测此类内容仍是一项重大挑战。因此，大多数现有的语义级检测方法难以可靠地识别它们。虽然某些基于LLM的检测方法显示出前景，但它们计算成本高，且受底层语言模型性能的限制，使其难以大规模部署。在这项工作中，我们旨在解决这一问题。我们提出了Paladin，它使用各种插入策略将触发器-标签关联嵌入到基础LLM中，将其改造为检测型LLM。当检测型LLM生成与钓鱼相关的内容时，它会自动包含可检测的标签，从而实现更轻松的识别。基于隐式和显式触发器与标签的设计，我们考虑了四种不同的场景。我们从隐蔽性、有效性和稳健性三个关键角度评估我们的方法，并与现有的基线方法进行比较。实验结果表明，我们的方法优于基线方法，在所有场景中均实现了超过90%的检测准确率。</span></span></p><p cid="n683" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s2522-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s2522-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n685" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">166、Pallas and Aegis: Rollback Resilience in TEE-Aided Blockchain Consensus</span></span></p><p cid="n686" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">许多拜占庭容错(BFT)共识算法利用可信组件来提高弹性并减少通信开销。然而，最近的研究发现，当可信组件崩溃、丢失状态或被克隆时，存在一个关键的回滚攻击漏洞。现有的防御方法要么将崩溃的副本视为拜占antine节点，从而增加副本数量，要么在组件间复制可信状态，这会带来巨大的性能开销，并且仅提供有限的容错能力。我们提出了一种稳健的替代方案：一种针对可信组件的安全状态保存机制，消除了在副本间复制可信状态的高昂成本。其核心是Aegis，这是首个专为使用可信组件的BFT协议设计的高效视图同步器。Aegis确保每个副本在任何视图中只有一个可信组件实例可以投票，即使可信组件在崩溃后重新启动或被敌对者克隆。在Aegis的基础上，我们引入了Pallas，这是首个能够在强敌对者控制固定数量的拜占庭副本并可能导致数量不定的可信组件崩溃的情况下保持安全性的BFT共识协议。我们确定了在部分同步条件下Pallas确保活跃性的敌对条件。在Amazon AWS上进行的大量地理分布式评估表明，Pallas在稳定条件下提供高性能且开销可忽略，吞吐量比现有协议高41%，延迟低29%。更重要的是，在其他协议失败的敌对条件下，Pallas仍能保持活跃性和优雅降级。</span></span></p><p cid="n687" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2443-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2443-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n689" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">167、Pando: Extremely Scalable BFT Based on Committee Sampling</span></span></p><p cid="n690" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">拜占庭容错（BFT）协议一直面临着可扩展性问题。事实上，随着副本数量n的增加，其性能会急剧下降。尽管已有大量工作试图实现可扩展性目标，但这些工作最多只能扩展到大约一百个副本，尤其是在低端机器上。在本文中，我们基于所谓的委员会采样方法开发了BFT协议，该方法选择一个小型委员会进行共识，并将结果传达给所有副本。然而，这种方法一直专注于拜占庭协议（BA）问题（仅考虑副本），而非拜占庭容错（BFT）问题（在客户端-副本模型中）；此外，该方法主要仅具有理论意义，因为实际上它适用于不切实际的大n值。我们基于委员会采样方法，在部分同步环境中构建了一个名为Pando的极其高效、可扩展且具有自适应安全性的BFT协议。我们在Amazon EC2上的评估表明，与现有协议相比，Pando可以轻松扩展到WAN环境中的一千个副本，实现62.57 ktx/秒的吞吐量。</span></span></p><p cid="n691" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s273-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s273-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n693" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">168、PANDORA: Lightweight Adversarial Defense for Edge IoT using Uncertainty-Aware Metric Learning</span></span></p><p cid="n695" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">资源受限的物联网设备的快速增长显著扩大了攻击面，暴露了网络中的关键漏洞。因此，依赖静态、基于签名的传统入侵检测系统已日益过时。现代攻击者现在采用复杂、自动化且通常是新颖（零日）的攻击，这些攻击可以轻易绕过此类传统防御。此外，现有基于机器学习的入侵检测模型在实际场景中往往难以处理概念漂移和无法泛化到未知威胁等挑战。为解决这些差距，我们引入了PANDORA（资源受限架构上的概率网络防御），这是一个用于检测边缘设备上零日攻击的新型端到端框架。PANDORA做出三项关键贡献：1）它学习不确定性感知的概率嵌入，以创建网络流量的鲁棒表示；2）它引入了一种新颖的概率流形结构和距离（PMSD）损失函数，实现了有效的零样本泛化；3）它利用高效的Mamba-专家混合（MoE）架构进行设备端部署。为验证我们的方法，我们还引入了TTDFIOTIDS2025数据集，这是一个新的高保真基准，包含复杂、程序生成的攻击。我们的广泛评估表明，PANDORA显著优于最先进的模型，在CICIDS2017上仅通过10次样本适应即可达到0.971的F1分数。关键的是，在域偏移条件下，其零样本检测准确率高达99%，并且在部署到树莓派时，保持约24 MB的低内存占用和高达4.26流/秒的吞吐量，证明了其在实时边缘安全中的实际可行性。</span></span></p><p cid="n696" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f713-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f713-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n698" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">169、Passive Multi-Target GUTI Identification via Visual-RF Correlation in LTE Networks</span></span></p><p cid="n699" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">LTE网络采用全球唯一临时标识符（GUTI）来保护用户免受永久国际移动用户身份（IMSI）的暴露，但我们表明，这些标识符可以通过被动观察在无需预先了解目标的情况下被解析并链接到特定设备。我们将带有时间戳的设备使用视觉观察与使用商用软件定义无线电（SDR）捕获的空中控制平面消息相关联。有限状态机（FSM）算法处理同步流以解析相机视场（FoV）内每个设备的GUTI，只要捕获相应的控制平面消息，仅需观察三次用户交互即可完成。在多个商业长期演进（LTE）网络进行的实地实验验证了多目标解析能力：在某些部署中，我们观察到GUTI可保持长达33天，且重新分配行为通常可被链接。一旦链接，这些长期存在的标识符通过被动监测寻呼消息和无线资源控制（RRC）消息，实现了从小区到寻呼区域范围的分层位置跟踪。与需要预先存在的标识符（如电话号码）和主动探测的主动IMSI捕获器或先前的GUTI攻击不同，我们的方法是仅监听模式，并可扩展到视场内的多个设备。</span></span></p><p cid="n700" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2487-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2487-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n702" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">170、PathProb: Probabilistic Inference and Path Scoring for Enhanced and Flexible BGP Route Leak Detection</span></span></p><p cid="n703" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">边界网关协议（BGP）缺乏内在安全性，使互联网容易受到路由泄漏等严重威胁。现有的检测方法存在二元分类僵化、误报率高以及权威AS关系数据稀疏等局限性。为应对这些挑战，本文提出了PathProb——一种新颖的范式，通过计算AS链路的拓扑感知概率分布和计算AS路径的合法性分数，灵活识别路由泄漏。我们的方法将蒙特卡洛方法与路由策略的整数线性规划公式相结合，以高效推导这些解决方案。我们使用真实的BGP路由跟踪和路由泄漏事件对PathProb进行了全面评估。结果表明，我们的推理模型在具有高置信度的验证数据集上优于最先进的方法。PathProb以98.45%的召回率检测真实世界的路由泄漏，同时将误报率比现有替代方法降低4.29~20.08个百分点。此外，PathProb的路径合法性评分使网络管理员能够动态调整路由泄漏检测阈值——根据其特定的误报容忍度和安全需求定制安全态势。最后，PathProb与新兴的路由缓解机制（如自治系统提供者授权（ASPA））无缝兼容，能够灵活集成以增强泄漏检测能力。</span></span></p><p cid="n704" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1691-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1691-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n706" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">171、Peering Inside the Black-Box: Long-Range and Scalable Model Architecture Snooping via GPU Electromagnetic Side-Channel</span></span></p><p cid="n707" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着深度神经网络（DNN）在自动驾驶和人脸识别等安全关键应用中的日益普及，它们也成为对抗性攻击的目标。然而，DNN的机密信息（包括模型架构）通常对攻击者隐藏。因此，对抗性攻击通常在黑盒环境下发起，这限制了其有效性。在本文中，我们提出了ModelSpy，一种基于GPU电磁（EM）泄漏的隐蔽DNN架构窥探攻击。ModelSpy能够在几米外甚至穿透墙壁提取完整的架构信息。ModelSpy基于一个关键观察：在DNN推理过程中，GPU会发出远场电磁信号，这些信号表现出特定于架构的幅度调制。我们开发了一个分层重建模型，从嘈杂的电磁信号中恢复细粒度的架构细节。为了提高对不同且不断演变的架构的可扩展性，我们设计了一个迁移学习方案，利用外部电磁泄漏与内部GPU活动之间的相关性。我们设计并实现了一个概念验证系统，以证明ModelSpy的可行性。我们在五款高端消费级GPU上的评估显示，ModelSpy在架构重建方面具有高准确性，包括97.6%的层分割准确率和94.0%的超参数估计准确率，工作距离可达6米。此外，ModelSpy重建的DNN与受害架构具有相当的性能，并能有效增强黑盒对抗性攻击。</span></span></p><p cid="n708" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s141-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s141-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n710" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">172、PhantomMap: GPU-Assisted Kernel Exploitation</span></span></p><p cid="n711" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">图形处理单元（GPU）已成为现代计算中的关键组件，推动了高性能渲染和并行处理的发展。其中，Arm公司的Mali GPU是在移动设备中部署最广泛的。与CPU端成熟且稳健的防御措施相比，GPU的安全保护仍然不足。因此，GPU已成为攻击者绕过CPU防御的首选目标。例如&#34;三角行动&#34;（Operation Triangulation）等重大事件已经证明，GPU端的漏洞可以被利用来危害系统安全。尽管威胁日益增加，对Mali GPU的全面深入的安全分析仍然缺失。为填补这一空白，我们首次对Mali GPU的内存映射机制进行了深入的安全分析，发现了两个新的安全弱点：分配-映射解耦和物理地址验证缺失。利用这些弱点，我们提出了PhantomMap，一种新颖的GPU辅助利用技术，可将有限的堆漏洞转化为强大的物理内存读写原语，无需特权能力或信息泄露即可绕过主流内核防御。为评估其安全影响，我们开发了一个静态分析工具，能够系统识别所有易受攻击的映射路径，在两种Mali驱动架构中发现了15个利用链。我们基于真实世界的CVE漏洞开发了15个端到端漏洞利用程序，进一步证明了PhantomMap的实用性，其中包括CVE-2025-21836的首个公开漏洞利用程序。最后，我们设计并实现了一个轻量级的驱动内缓解措施，在Pixel 6和Pixel 7设备上以最小的性能开销消除了根本原因。</span></span></p><p cid="n712" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f201-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f201-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n714" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">173、PhantomMotion: Laser-Based Motion Injection Attacks on Wireless Security Surveillance Systems</span></span></p><p cid="n715" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">无线安全监控系统因其日益提高的性价比而被广泛部署。运动检测通常被集成到这些系统中，作为其安全功能的核心，用于检测是否有人在监控范围内，然后触发系统开始录制或通知财产所有者。在本文中，我们提出了PhantomMotion，一种新的攻击框架，用于欺骗这些安全系统的运动检测功能。它可以通过将激光束瞄准运动检测范围来秘密地创建虚假运动刺激，并通过嗅探无线流量确认系统对刺激的响应。PhantomMotion不需要任何专业设备，也不需要在监控区域内进行物理运动。它包含一个集成了激光控制和WiFi嗅探的新型硬件平台，以及一种新的运动注入生成机制。我们开发了一款智能手机应用程序来实现PhantomMotion，并在18种流行的无线运动激活安全系统上验证了其有效性。实验结果表明，PhantomMotion能够始终生成虚假运动来成功触发这些系统，平均耗时12.8秒，激光点移动平均距离为1.1米。值得注意的是，我们验证了PhantomMotion可以在高达120米的距离上有效工作。</span></span></p><p cid="n716" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1454-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1454-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n718" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">174、Phishing in Wonderland: Evaluating Learning-Based Ethereum Phishing Transaction Detection and Pitfalls</span></span></p><p cid="n719" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">钓鱼攻击仍然是以太坊生态系统的一个重大威胁，占以太坊相关网络犯罪的50%以上，并促使基于机器学习的防御措施兴起。本文提出了一个综合框架，通过解决特征选择、类别不平衡、模型鲁棒性和算法优化等关键挑战，来增强以太坊交易中的钓鱼检测。通过对现有方法的系统性评估，我们确定了实践中的主要差距，特别是在特征处理和不可持续的性能提升方面。我们的分析和实证评估表明，所提出的框架提高了检测的泛化能力和有效性。这些研究结果强调了需要完善检测策略，以应对区块链领域日益复杂的钓鱼战术。</span></span></p><p cid="n720" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f694-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f694-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n722" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">175、PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT</span></span></p><p cid="n723" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们提出了PhishLang，这是首个完全基于客户端的反钓鱼框架，以基于Chromium的浏览器扩展形式实现。PhishLang利用轻量级语言模型(MobileBERT)实现钓鱼网站的实时、设备端检测。与难以应对规避性威胁的传统启发式或静态特征模型，以及对于客户端使用而言资源消耗过大的深度学习方法不同，PhishLang分析页面源代码的上下文结构，在检测性能上与几种最先进的模型相当，同时内存消耗比类似架构低多达7倍。在为期3.5个月的期间内，我们实时部署了该框架，成功识别了约26,000个钓鱼URL，其中许多是流行反钓鱼黑名单未检测到的，从而证明了PhishLang辅助当前检测措施的潜力。另一方面，该浏览器扩展超越了多种反钓鱼工具，在零日攻击期间检测到超过91%的威胁。PhishLang还表现出强大的对抗鲁棒性，通过解析器级防御和对抗性重训相结合的方式，抵抗了16类真实问题空间的规避攻击。为了帮助终端用户和研究社区，我们已经开源了PhishLang框架和浏览器扩展。</span></span></p><p cid="n724" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1037-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1037-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n726" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">176、PhyFuzz: Detecting Sensor Vulnerabilities with Physical Signal Fuzzing</span></span></p><p cid="n727" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">传感器漏洞可被物理信号攻击利用，导致传感器测量错误，危及依赖传感器做出关键决策的系统。尽管已有数百项研究发现了众多传感器漏洞，但它们都依赖于人工专家分析，需要耗时耗力的试错过程。缺乏自动化方法辅助检测传感器漏洞，已成为连接传感器安全研究与工业应用之间鸿沟的主要障碍。本文提出PhyFuzz，一种新的物理信号模糊测试范式，它依赖物理测试信号来检测现有及潜在的新型传感器漏洞，无需人工干预。为应对物理信号模糊测试带来的前所未有的挑战，如信号参数的无限搜索空间和多样化传感器硬件的黑盒设计，我们设计了一种独特的模糊测试算法，能够高效构建测试信号，并对传感器漏洞识别和评估进行有效的特征离散化。我们实现了PhyFuzz原型系统，支持声学、激光和电磁信号的模糊测试。实验表明，该系统能够在9种不同类型的13个传感器上识别出46个漏洞，其中包括6个未公开的案例。</span></span></p><p cid="n728" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f29-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f29-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n730" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">177、PIRANHAS: PrIvacy-Preserving Remote Attestation in Non-Hierarchical Asynchronous Swarms</span></span></p><p cid="n731" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">远程认证是评估远程设备完整性的基本安全机制。在实践中，现有协议缺乏公开可验证性和交互要求，这阻碍了认证方案的广泛采用。Ebrahimi 等人 (NDSS&#39;24) 的最新工作构建了公开可验证的非交互式远程认证，但却忽略了认证敏感系统的另一个重要要求：隐私保护。在物联网集群中，许多设备可能处理敏感数据，这些设备应产生单一的认证证明，同样存在此类需求。在本文中，我们同时应对这两个挑战。我们提出了 PIRANHAS，一种针对单个设备和集群的公开可验证、异步且匿名的认证方案。我们利用 zk-SNARKs 将任何经典的对称远程认证方案转换为非交互式、公开可验证且匿名的方案。验证者仅确认认证的有效性，而不了解任何关于相关设备的识别信息。对于物联网集群，PIRANHAS 使用递归 zk-SNARKs 对整个集群的认证证明进行聚合。我们的系统支持任意网络拓扑结构，并允许节点动态加入和离开网络。我们为单设备和集群场景提供了形式化安全证明，表明我们的构造满足所需的安全保证。此外，我们使用 Noir 和 Plonky2 框架提供了我们方案的开源实现，实现了仅 356ms 的聚合运行时间。</span></span></p><p cid="n732" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f526-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f526-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n734" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">178、Pitfalls for Security Isolation in Multi-CPU Systems</span></span></p><p cid="n736" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在嵌入式系统中，将多个CPU集成到单个系统芯片(SoC)中可以实现更高的性能，并将任务分离为独立的固件和优化架构。例如，ARM Cortex-M4核心可以运行主固件，而Cortex-M0核心可以运行实时操作系统(RTOS)。此类集成的安全影响仍不明确，例如，如果一个攻击者在某个CPU上执行代码，是否能够完全攻破第二个CPU或泄露受保护数据。在这项工作中，我们系统地识别了此类集成导致的安全问题，特别是与内存和外设访问控制相关的问题。这些问题源于在新多CPU系统中重用单CPU安全机制，如内存保护单元(MPU)。我们确定了此类系统中可能存在的四种主要攻击向量，并发现市场上大量系统似乎存在漏洞。这些攻击向量可能导致对另一个CPU受保护内存的任意读写，甚至导致代码执行。此外，我们发现一种流行的开源RTOS FreeRTOS[17]的通信机制（被建议作为多CPU系统上固件间的通信机制）在多CPU场景中引入了代码执行漏洞。随后，我们通过实施四种攻击向量验证了我们的理论预测，并证明了其实际有效性。此外，我们发现在一个案例中，发现的攻击面可能导致自定义可信执行环境(TEE)实现的攻破。我们向供应商负责任地披露了我们的发现，导致发布了安全公告并对专有网络栈实现进行了修复。</span></span></p><p cid="n737" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f971-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f971-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n739" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">179、PortRush: Detect Write Port Contention Side-Channel Vulnerabilities via Hardware Fuzzing</span></span></p><p cid="n740" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">CPU漏洞在现代CPU架构中持续构成安全挑战。在CPU漏洞中，写端口竞争——由多个功能模块同时竞争有限的共享写端口引起——仍未得到充分研究。本文研究了CPU中的写端口竞争侧信道漏洞，并提出了</span></span><span md-inline="strong" style="box-sizing: border-box;"><strong style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">PortRush</span></span></strong></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">，一种新型模糊测试框架，用于在寄存器传输级（RTL）检测和验证此类漏洞。首先，PortRush构建</span></span><span md-inline="strong" style="box-sizing: border-box;"><strong style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">写请求图（WRG）</span></span></strong></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">，通过建模目标共享存储单元的功能模块之间的写路径和优先关系，静态识别潜在的写端口竞争实例。其次，在WRG中，PortRush实现了</span></span><span md-inline="strong" style="box-sizing: border-box;"><strong style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">分层聚合和解码</span></span></strong></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">方法，通过监控设计层次结构中的相关硬件信号，高效检测写端口竞争。第三，PortRush采用</span></span><span md-inline="strong" style="box-sizing: border-box;"><strong style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">竞争引导的硬件模糊测试</span></span></strong></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">方法，触发写端口竞争，并将竞争触发的指令序列与瞬时执行攻击模式自动结合，从而验证写端口竞争侧信道漏洞。我们在三个RISC-V CPU（BOOM、NutShell和Rocket Core）上评估了PortRush，证明了其在识别和触发写端口竞争方面的有效性。此外，我们验证了所发现的漏洞可在实际的写端口竞争攻击场景中被利用。基于这些漏洞，我们提出了两种新型攻击向量：</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Birgus变体</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">，利用重排序缓冲区中物理寄存器文件的竞争；以及</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">MSHRush</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">，利用L1数据缓存中加载/存储单元（LSU）与缺失状态处理寄存器（MSHR）之间的竞争，以诱导依赖于秘密的执行延迟。我们还为CPU开发者提出了缓解此类漏洞的策略。</span></span></p><p cid="n741" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f587-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f587-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n743" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">180、Practical Traceable Over-Threshold Multi-Party Private Set Intersection</span></span></p><p cid="n744" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">带阈值的多方私密集合交集（MP-PSI）通过披露至少出现在t个参与者集合中的元素，而非要求元素出现在所有n个集合中，从而增强了MP-PSI的灵活性。在每位参与者负责其数据集的场景中，例如数字取证，带阈值的MP-PSI应披露交集元素及其对应持有者，以便元素可追溯，从而保证交集的可靠性。我们将支持可追溯性的带阈值MP-PSI称为可追溯超阈值多方私密集合交集（T-OT-MP-PSI）。然而，此类协议的研究仍然有限，当前解决方案能够抵抗t-2个半诚实参与者，但代价是巨大的计算开销。在本文中，我们提出了两种新颖的可追溯OT-MP-PSI协议。第一种是高效可追溯OT-MP-PSI（ET-OT-MP-PSI），它将Shamir秘密共享与可忽略可编程伪随机函数相结合，在抵抗最多t-2个半诚实参与者的同时显著提高了效率。第二种是增强安全性的可追溯OT-MP-PSI（ST-OT-MP-PSI），它通过进一步利用可忽略线性评估协议，实现了抵抗多达n-1个半诚实参与者的安全性。与Mahdavi等人最近的Traceable OT-MP-PSI协议相比，我们的协议消除了某些特殊参与者不共谋的安全假设，并提供了更强的安全保证。我们实现了所提出的协议并在各种设置下进行了广泛实验。我们将我们的协议与Mahdavi等人的协议进行了性能比较。尽管我们的可追溯OT-MP-PSI协议增强了安全性，但实验结果表明其具有高效率。例如，给定5个参与者，阈值为3，集合大小为2^14时，我们的ET-OT-MP-PSI协议比Mahdavi等人的协议快15056倍，而ST-OT-MP-PSI协议快505倍。</span></span></p><p cid="n745" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s38-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s38-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n747" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">181、PriSrv+: Privacy and Usability-Enhanced Wireless Service Discovery with Fast and Expressive Matchmaking Encryption</span></span></p><p cid="n748" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">服务发现是无线网络中的基本过程，使设备能够动态地查找并与服务通信，对于5G和物联网等现代系统的无缝运行至关重要。本文介绍了PriSrv+，一种针对现代无线网络和资源受限环境的先进隐私和可用性增强型服务发现协议。PriSrv+基于PriSrv（NDSS&#39;24），通过解决在表达性、隐私性、可扩展性和效率方面的关键局限性，同时保持与广泛使用的无线协议（如mDNS、BLE和Wi-Fi）的兼容性。PriSrv+的一个关键创新是开发了快速且表达性强的匹配加密（FEME），这是第一个能够支持具有无界属性宇宙的表达性访问控制策略的匹配加密方案，允许使用任意字符串作为属性。FEME显著增强了服务发现的灵活性，同时确保了强大的消息和属性隐私。与PriSrv相比，PriSrv+优化了加密操作，加密速度提高了7.62倍，解密速度提高了6.23倍，并将密文大小减少了87.33%。此外，与PriSrv相比，PriSrv+将服务广播的通信成本降低了87.33%，匿名相互认证的通信成本降低了86.64%。形式化安全证明确认了FEME和PriSrv+的安全性。在多个平台上的广泛评估表明，与现有最先进的协议相比，PriSrv+实现了卓越的性能、可扩展性和效率。</span></span></p><p cid="n749" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s87-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s87-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n751" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">182、PrivATE: Differentially Private Average Treatment Effect Estimation for Observational Data</span></span></p><p cid="n752" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">因果推断在多学科科学研究中发挥着关键作用。从观测数据中估计因果效应，特别是平均处理效应（ATE），已引起广泛关注。然而，从现实世界观测数据计算ATE会对用户造成严重的隐私风险。差分隐私提供了严格的理论保证，已成为隐私保护数据分析的标准方法。然而，现有的差分隐私ATE估计研究依赖于特定假设，提供的隐私保护有限，或无法提供全面的信息保护。为此，我们引入了PrivATE，一个确保差分隐私的实用ATE估计框架。实际上，不同场景需要不同程度的隐私保护。例如，在教育评估中，只有测试成绩通常是敏感信息，而所有类型的医疗记录数据通常都是私有的。为了适应不同的隐私需求，我们在PrivATE中设计了两个级别的隐私保护（即标签级和样本级）。通过推导自适应匹配限制，PrivATE有效平衡了噪声引起的误差和匹配误差，从而实现了更准确的ATE估计。我们的评估验证了PrivATE的有效性。在所有数据集和隐私预算下，PrivATE均优于基线方法。</span></span></p><p cid="n753" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1350-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1350-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n755" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">183、PrivCode: When Code Generation Meets Differential Privacy</span></span></p><p cid="n756" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）在代码生成和补全方面表现出色。然而，在私有数据集上对这些模型进行微调可能会引发隐私和专有性问题，例如敏感个人信息的泄露。差分私有（DP）代码生成通过生成保留统计特性同时减少隐私泄露担忧的合成数据集，为保护敏感代码提供了理论保证。然而，DP代码生成面临着严格的语法依赖性和隐私-效用权衡的显著挑战。我们提出了PrivCode，这是首个专门为代码数据集设计的DP合成器。它采用两阶段框架来提高隐私性和效用性。在第一阶段，称为&#34;隐私净化&#34;，PrivCode通过使用DP-SGD训练模型并引入语法信息来保留代码结构，生成符合DP要求的合成代码。第二阶段，称为&#34;效用提升&#34;，在无隐私风险的合成代码上对更大的预训练LLM进行微调，以减轻DP造成的效用损失，提高生成代码的效用性。在四个LLMs上的广泛实验表明，在四个基准测试的各种测试任务中，PrivCode生成的代码具有更高的效用。实验还证实了它在不同隐私预算下保护敏感数据的能力。我们在匿名链接提供了复制包。</span></span></p><p cid="n757" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f936-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f936-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n759" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">184、PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning</span></span></p><p cid="n760" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">最近，离线强化学习（RL）已成为一种流行的强化学习范式。在离线强化学习中，数据提供者共享预先收集的数据集——无论是作为单个转换还是形成轨迹的转换序列——以实现强化学习模型（也称为智能体）的训练，而无需直接与环境交互。与传统强化学习相比，离线强化学习减少了与环境的交互，并在导航任务等关键领域已证明其有效性。同时，关于离线强化学习数据集隐私泄露的担忧也随之出现。为了保护离线强化学习数据集中的私人信息，我们提出了首个差分隐私（DP）离线数据集合成方法PrivORL，该方法分别利用扩散模型和扩散转换器在差分隐私条件下合成转换和轨迹。然后，合成数据集可以安全地发布用于下游分析和研究。PrivORL采用在公共数据集上预训练合成器，然后使用差分随机梯度下降（DP-SGD）在敏感数据集上进行微调的流行方法。此外，PrivORL引入了由好奇心驱动的预训练，该方法利用好奇心模块的反馈来多样化合成数据集，从而能够生成与敏感数据集高度相似的多样化合成转换和轨迹。在五个敏感离线强化学习数据集上的广泛实验表明，与基线方法相比，我们的方法在差分隐私转换和轨迹合成中实现了更好的效用和保真度。复制包可通过匿名链接获取。</span></span></p><p cid="n761" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f149-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f149-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n763" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">185、Prompt Injection Attack to Tool Selection in LLM Agents</span></span></p><p cid="n764" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">工具选择是LLM智能体的关键组成部分。一种流行的方法遵循两步过程——检索和选择——从工具库中为给定任务选择最合适的工具。在这项工作中，我们引入了ToolHijacker，这是一种针对无框场景中工具选择的新型提示注入攻击。ToolHijacker将恶意工具文档注入工具库，以操纵LLM智能体的工具选择过程，迫使其始终为攻击者选择的目标任务选择攻击者的恶意工具。具体而言，我们将此类工具文档的制定表述为一个优化问题，并提出了一种两阶段优化策略来解决它。我们广泛的实验评估表明，ToolHijacker非常有效，在应用于工具选择时，显著优于现有的基于手动和自动化的提示注入攻击。此外，我们探索了各种防御措施，包括基于预防的防御（StruQ和SecAlign）和基于检测的防御（已知答案检测、DataSentinel、困惑度检测和窗口化困惑度检测）。我们的实验结果表明，这些防御措施不足，凸显了开发新防御策略的迫切需求。</span></span></p><p cid="n765" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s675-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s675-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n767" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">186、ProtocolGuard: Detecting Protocol Non-compliance Bugs via LLM-guided Static Analysis and Dynamic Verification</span></span></p><p cid="n769" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">网络协议实现应严格遵循其规范以确保可靠和安全的通信。然而，自然语言规范的固有歧义常导致开发者的误解，使协议实现偏离标准行为。这些偏差会导致细微的不合规错误，引发互操作性和关键安全问题。与内存损坏错误不同，这类错误通常不表现出明显的错误行为，导致现有的错误预言机制不足以全面检测它们。此外，现有工作需要大量手动工作来验证发现和分析根本原因，严重限制了它们的实际可扩展性。</span></span></p><p cid="n770" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">本文提出了ProtocolGuard，一个新颖的框架，通过结合大语言模型（LLM）引导的静态分析与基于模糊测试的动态验证，系统性地检测不合规错误。ProtocolGuard首先使用混合方法从协议规范中提取规范性规则，并执行LLM引导的程序切片，提取与每条规则相关的代码片段。然后，它利用LLM检测这些规则与代码逻辑之间的语义不一致，并动态验证这些错误是否可以被触发。为便于错误验证，ProtocolGuard首先使用LLM自动生成断言语句并对代码进行插桩，将静默的不一致转变为可观察的断言失败。接着，借助LLM生成更有可能触发错误的初始测试用例进行动态验证。最后，ProtocolGuard动态测试插桩后的代码，确认错误识别并生成概念验证测试用例。</span></span></p><p cid="n771" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们实现了ProtocolGuard的原型，并在11个广泛使用的协议实现上对其进行了评估。ProtocolGuard以高精度成功发现了158个不合规错误，其中70个已得到确认，且大多数可以转换为断言并进行动态验证。与现有最先进工具的对比表明，在错误检测能力方面，ProtocolGuard在精确率和召回率上都优于它们。</span></span></p><p cid="n772" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f521-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f521-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n774" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">187、Pruning the Tree: Rethinking RPKI Architecture from the Ground up</span></span></p><p cid="n775" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">资源公钥基础设施（RPKI）是BGP的关键安全机制，但随着其采用规模的扩大，其架构复杂性日益成为关注点。当前RPKI设计大量重用了传统PKI组件，如X.509 EE证书、ASN.1编码和基于XML的存储库协议，这些引入了过多的密码验证、冗余元数据以及在存储和处理方面的低效。我们表明，这些设计选择虽然基于既定标准，但造成了显著的性能瓶颈，增加了攻击面，并阻碍了大规模互联网部署的可扩展性。在本文中，我们首次对RPKI设计中复杂性的根本原因进行了系统性分析，并通过实验量化了它们在现实世界中的影响。我们表明，RPKI依赖方超过70%的验证时间花费在证书解析和签名验证上，其中大部分是不必要的。基于这一见解，我们引入了改进的RPKI（iRPKI），这是一种向后兼容的重新设计，在保留所有安全保证的同时显著减少了协议开销。iRPKI消除了EE证书和ROA签名，合并了撤销和完整性对象，用Protobuf替换了冗长的编码，并重新构造了存储库元数据以实现更高效的访问。我们通过实验证明，在Routinator验证器中实现的iRPKI实现了处理时间20倍的加速，带宽需求18倍的改进，缓存内存占用8倍的减少，同时消除了已在RPKI软件中导致至少10个漏洞的漏洞类别。iRPKI显著提高了在互联网中特别是在受限环境中大规模部署RPKI的可行性。我们的设计可以增量部署而不会影响现有操作。我们开源了我们的设计、对象模板、发布点软件和RP实现，以促进iRPKI集成到当前RPKI部署中，并能够复现我们的研究。我们进一步提供了如何从我们提出的改进中推导新RPKI规范的建议，以促进标准化。</span></span></p><p cid="n776" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s823-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s823-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n778" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">188、Prεεmpt: Sanitizing Sensitive Prompts for LLMs</span></span></p><p cid="n780" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型(LLMs)的兴起带来了新的隐私挑战，特别是在推理过程中，提示中的敏感信息可能暴露给专有的LLM API。在本文中，我们解决了在保持响应质量的同时正式保护提示中包含的敏感信息的问题。为此，首先，我们引入了一种受密码学启发的&#34;提示净化器&#34;概念，用于转换输入提示以保护其敏感标记。其次，我们提出了Pr$epsilonepsilon$mpt，一个实现提示净化器的系统，专注于仅能从各个标记中推导出的敏感信息。Pr$epsilonepsilon$mpt将敏感标记分为两类：(1)LLM的响应仅依赖于格式的标记(如社会保障号、信用卡号)，对此我们使用格式保留加密(FPE)；(2)响应依赖于特定值的标记(如年龄、薪资)，对此我们应用度量差分隐私(mDP)。我们的评估表明，Pr$epsilonepsilon$mpt是一种实现有意义隐私保证的实用方法，与未净化的提示相比保持了高效用，并优于先前的方法。</span></span></p><p cid="n781" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1277-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1277-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n783" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">189、Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security</span></span></p><p cid="n784" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">多模态大型语言模型（MLLMs）在跨模态理解方面展现出令人印象深刻的能力，但尽管拥有强大的文本安全机制，仍易通过视觉输入受到对抗性攻击。这些弱点源于两个核心问题：视觉表示的连续性使得基于梯度的攻击成为可能，以及基于文本的安全机制无法充分迁移到视觉内容。我们提出了Q-MLLM，一种新颖的架构，通过集成两级向量量化来创建对抗性攻击的离散瓶颈，同时保留多模态推理能力。通过在像素块和语义级别对视觉表示进行离散化，Q-MLLM能够阻断攻击路径并弥合跨模态安全对齐的差距。我们的两阶段训练方法确保了稳健的学习同时保持模型效用。实验表明，Q-MLLM在抵御越狱攻击和有毒图像攻击方面的防御成功率显著优于现有方法。值得注意的是，除一个可争议的案例外，Q-MLLM对越狱攻击实现了完美的防御成功率（100%），同时在多个效用基准测试上保持有竞争力的性能，且推理开销最小。这项研究确立了向量量化作为安全多模态AI系统的有效防御机制，无需昂贵的特定安全微调或检测开销。</span></span></p><p cid="n785" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s407-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s407-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n787" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">190、QNBAD: Quantum Noise-induced Backdoor Attacks against Zero Noise Extrapolation</span></span></p><p cid="n788" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">变分量子算法（VQA）已成为在嘈杂中等规模量子（NISQ）时代实现实用量子优势的最有前景的范例之一。为了提高VQA在嘈杂硬件上的计算精度，零噪声外推（ZNE）已成为一种广泛采用且有效的错误缓解技术。然而，对ZNE的日益依赖也增加了识别潜在对抗性攻击的重要性。我们审视了现有的后门攻击，并强调了它们为何难以破坏ZNE。具体而言，仅修改电路结构的量子后门攻击只会移动理想输出而不影响噪声相关的外推过程，从而使ZNE保持完整。同样，不考虑设备特定噪声而训练的参数级后门在不同硬件平台上表现出不一致的行为，导致不可靠或无效的攻击。基于这些观察，我们发现了一类新的后门漏洞，专门针对ZNE的独特属性。在本研究中，我们提出了QNBAD，这是一种针对ZNE的新型隐蔽后门攻击。QNBAD经过精心设计，可在大多数设备上保持变分量子电路的正确功能。然而，在特定的噪声模型下，它利用量子噪声与电路结构之间的微妙相互作用，系统性地操纵不同噪声水平下的采样期望值。这种有针对性的干扰破坏了ZNE拟合过程，并导致显著偏差的最终估计。与先前的后门方法相比，QNBAD在四个平台和六个应用中实现了绝对误差放大1.68倍至11.7倍的显著提升。此外，它在各种拟合函数和ZNE变体中保持有效。</span></span></p><p cid="n789" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1665-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1665-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n791" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">191、ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits</span></span></p><p cid="n792" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">处理器设计依赖于迭代修改和重用成熟的设计。然而，这种对先前设计的重用也导致多个处理器之间存在相似的漏洞。随着处理器通过迭代修改变得越来越复杂，高效检测现代处理器中的漏洞变得至关重要。受软件模糊测试的启发，硬件模糊测试最近已证明其在检测处理器漏洞方面的有效性。然而，据我们所知，现有的处理器模糊测试器单独测试每个设计，缺乏理解先前处理器中已知漏洞的能力，无法微调模糊测试以识别相似或新的漏洞变体。为了解决这一差距，我们提出了</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">ReFuzz</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">，一个自适应模糊测试框架，它利用上下文老虎机来重用来自先前处理器的高度有效测试，以在给定ISA内测试目标处理器（PUT）。通过智能修改能触发先前处理器漏洞的测试，</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">ReFuzz</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">能够检测PUT中的相似漏洞和新变体。</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">ReFuzz</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">发现了三个新的安全漏洞和两个新的功能错误。</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">ReFuzz</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">通过重用触发先前处理器中已知漏洞的测试，检测到一个漏洞。一个功能错误存在于共享设计模块的三个处理器中。第二个错误有两个变体。此外，与现有的模糊测试器相比，</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">ReFuzz</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">通过重用高度有效的测试来提高覆盖率效率，实现了平均511.23倍的覆盖率加速和高达9.33%的额外总覆盖率。</span></span></p><p cid="n793" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f118-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f118-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n795" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">192、Repairing Trust in Domain Name Disputes Practices: Insights from a Quarter-Century’s Worth of Squabbles</span></span></p><p cid="n796" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">《统一域名争议解决政策》（UDRP）旨在平衡两个相互竞争的目标：赋予商标持有人迅速解决域名滥用的权力——例如销售经常绕过黑名单等技术保护措施的假冒商品，并保护注册人免受过度主张商标的当事人采取的激进法律策略。自实施以来，UDRP已成为超过一千二百个域名扩展的实际争议解决机制，比最初的三个有了显著增加。然而，尽管取得了成功，批评者认为该政策助长了破坏信任和公平的做法。不幸的是，由于缺乏大规模结构化数据，有意义的改革努力陷入停滞，这限制了实证评估，并使基础性问题在过去二十多年中一直悬而未决。为解决这一长期存在的空白，我们训练了模型从90,153个UDRP争议程序中提取结构化数据，从而实现了迄今为止对该政策最全面的实证分析。我们的研究结果揭示了几个问题，显示在几乎所有争议中近三分之一的案件存在&#34;法庭选购&#34;现象，43个案例中存在潜在的利益冲突，以及许多当事人的延迟回应时间远超预期——所有这些都影响了UDRP的感知公平性和效率。除了侵蚀信任外，这些问题还带来了严重的安全挑战：在专家组下令转移域名后，2,751个恶意域名仍在恶意行为者控制下长达四个月。总体而言，我们的研究结果强调了政策改革的必要性，以帮助恢复信任并提高互联网应对商标侵权的实际标准的透明度。基于我们的发现，我们建议引入更多自动化、加强监督和执行更明确的合规规则，以确保UDRP继续成为基于商标的名称争议的可靠工具——特别是在互联网随着新的通用顶级域名（2026年）和日益敌对的数字环境不断扩张的背景下。</span></span></p><p cid="n797" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s174-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s174-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n799" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">193、Rethinking Fake Speech Detection: A Generalized Framework Leveraging Spectrogram Magnitude</span></span></p><p cid="n800" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">受深度学习进步驱动的语音合成技术已取得了显著的逼真效果，使其能够在各个领域实现多样化应用。然而，这些技术也可能被用来生成虚假语音，带来重大风险。尽管现有的虚假语音检测方法在受控环境中表现出有效性，但它们往往难以推广到未见过的场景，包括新的合成模型、语言和录音条件。此外，许多现有方法依赖于特定假设，且缺乏对虚假语音中固有伪影的全面理解。本文通过提出一种专注于分析语谱图幅度的新视角，重新思考了虚假语音检测任务。通过广泛分析，我们发现合成语音在语谱图的幅度表示中始终表现出伪影，如纹理细节减少和不同幅度范围的不一致性。利用这些见解，我们引入了一种新颖的无假设且通用的虚假语音检测框架。该框架基于幅度将语谱图分层表示，并利用二维和三维表示在空间和离散余弦变换（DCT）域中检测伪影。这种设计使框架能够有效捕捉虚假语音中固有的细粒度伪影和合成不一致性。大量实验表明，该框架在几个广泛使用的公共音频深度伪造数据集上取得了最先进的性能。此外，在涉及黑盒网络语音克隆API的真实场景评估中，突显了该框架的鲁棒性和实际适用性， consistently优于基线方法。</span></span></p><p cid="n801" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1024-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1024-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n803" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">194、Revealing The Secret Power: How Algorithms Can Influence Content Visibility on Twitter/X</span></span></p><p cid="n804" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">近年来，社交网络推荐算法的不透明设计和公众对其的有限理解引发了人们对信息曝光可能被操纵的担忧。降低内容可见性，即所谓的&#34;影子封禁&#34;，可能有助于限制有害内容；然而，它也可能被用来压制不同声音。这促使我们需要更大的透明度和对这一做法的更好理解。在本文中，我们通过对两个Twitter/X数据集进行大规模定量分析来研究可见性变化的存在，这些数据集包含来自900多万用户的超过4000万条推文，重点关注围绕乌克兰-俄罗斯冲突和2024年美国总统大选的讨论。我们使用浏览量来检测可见性降低或增加的模式，并检查这些模式如何与用户观点、社会角色和叙事框架相关联。我们的分析表明，算法系统性地惩罚包含外部资源链接的推文，将其可见性降低多达8倍，而不管其意识形态立场或来源可靠性如何。相反，内容可见性可能会根据产生它的特定账户而被惩罚或青睐，正如比较基辅独立报和RT.com的推文或唐纳德·特朗普和卡玛拉·哈里斯的推文时所观察到的那样。总体而言，我们的工作强调了内容审核和推荐系统透明度的重要性，以保护公共话语的完整性并确保对在线平台的公平访问。</span></span></p><p cid="n805" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s718-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s718-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n807" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">195、Revisiting Differentially Private Hyper-parameter Tuning</span></span></p><p cid="n808" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们研究了差分隐私在超参数调优中的应用，该过程涉及从多个候选运行中选择最佳运行。与许多隐私学习算法（包括普遍使用的DP-SGD）不同，选择最佳运行的隐私影响常常被忽视。尽管最近的研究提出了针对调优过程的通用隐私选择解决方案，但一个悬而未决的问题仍然存在：这种隐私上界是否紧密？本文从实证和理论两方面探讨了这一问题。最初，我们提供的研究证实了当前隐私分析中关于隐私选择的结论在一般情况下确实是紧密的。然而，当我们具体研究白盒环境下的超参数调优问题时，这种紧密性便不再成立。这一点首先通过对调优过程进行隐私审计得到证明。我们的研究结果表明，即使在强大的审计设置下，当前的理论隐私边界与经验隐私泄露之间仍存在显著差距。这一差距促使我们进行后续的理论研究，由于超参数调优具有独特性质，我们为其提供了改进的隐私上界。我们的改进边界带来了更好的效用。与之前仅限于特定参数配置的分析相比，我们的分析还展示了更广泛的应用性。总体而言，我们对理解因&#34;选择&#34;导致的隐私退化做出了贡献。</span></span></p><p cid="n809" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s447-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s447-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n811" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">196、Robust Fraud Transaction Detection: A Two-Player Game Approach</span></span></p><p cid="n812" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于机器学习（ML）的欺诈检测系统被企业广泛采用，以减少欺诈活动造成的经济损失。然而，欺诈者具有智能性且快速演变，采用先进技术伪造交易特征以规避检测系统。更糟糕的是，由于这些伪造过程不受小范围限制，基于小规模扰动的现有鲁棒性增强方法无效。检测不受限制扰动的欺诈活动显著增加了欺诈检测的不确定性，这仍然是一个开放性问题。为解决这一问题，我们提出了GAMER，一个基于双人博弈的鲁棒欺诈检测系统，在检测欺诈活动时实现了高准确性和强鲁棒性。具体而言，GAMER利用特征选择主动对抗欺诈检测中的智能欺诈者（即选择较少的特征以减少特征伪造的组合），并创新地将检测过程表述为双人博弈。通过求解双人博弈的均衡点，GAMER计算特征选择的最优概率，该概率考虑了欺诈者所有可能的伪造策略。基于均衡点的选择概率不仅最小化了欺诈者获得的收益，从而阻止他们发起伪造；还使系统能够在检测欺诈活动时选择鲁棒特征（即不太可能被伪造的特征），增强了系统在欺诈检测中的鲁棒性。我们的理论和实验结果验证了威慑和鲁棒性增强的特性。此外，对全球领先在线支付企业遭受的真实攻击进行的实验表明，GAMER优于传统的鲁棒性增强技术，在为期两个月的欺诈检测中平均将F1分数提高了67.5%。</span></span></p><p cid="n813" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1611-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1611-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n815" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">197、ropbot: Reimaging Code Reuse Attack Synthesis</span></span></p><p cid="n817" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">代码重用攻击是现代基于内存损坏攻击的最关键基石之一。然而，将代码片段(gadgets)拼接在一起的任务仍然是一个耗时且手动的过程。过去十年间已发表大量研究旨在自动化解决这个问题，但实践中很少被采用。这些解决方案在性能或支持的架构方面通常不切实际，或者无法生成有效的代码链。系统分析表明，它们都采用生成-测试方法，即首先枚举所有代码片段，然后使用符号执行或SMT求解器来推理哪些代码片段可以组合成链。不幸的是，这种方法随可用代码片段的数量呈指数级扩展，从而限制了在较大二进制文件上的可扩展性。在这项工作中，我们重新审视这一基本策略，并提出了一种新的代码片段分组方法，称为ROPBlock，它与代码片段有一个关键区别：ROPBlock保证可以链接。我们将ROPBlock的概念与图搜索算法相结合，提出了一种代码链接方法，与先前的工作相比显著提高了性能。我们将设置寄存器为攻击者指定值的时间复杂度从O(2^n)降低到O(n)。这在实践中带来了2-3个数量级的加速。同时，ROPBlock使我们能够建模复杂的代码片段——例如涉及ret2csu或包含条件分支的代码片段——而大多数其他方法在设计上无法考虑这些。由于ROPBlock与架构无关，我们的方法可以应用于多种架构。我们的原型工具ropbot在评估的所有37个二进制文件上平均2.5秒内即可生成调用dup-dup-execve的复杂真实世界代码链。除了一种方法外，所有其他方法都无法在此场景下生成任何代码链。对于需要设置六个寄存器值的困难场景——mmap链，ropbot找到的目标数量是第二佳技术的5倍。为了展示其多功能性，我们在x64、MIPS、ARM和AArch64上评估了ropbot。我们仅通过添加十二行代码就在不到两小时内添加了RISC-V支持。最后，我们证明ropbot在各自的数据集上优于所有现有工具。</span></span></p><p cid="n818" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f845-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f845-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n820" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">198、Rounding-Guided Backdoor Injection in Deep Learning Model Quantization</span></span></p><p cid="n821" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">模型量化是将深度学习模型部署在资源受限环境中的常用技术。然而，它也可能引入先前被忽视的安全风险。在这项工作中，我们提出了QuRA，一种利用模型量化来嵌入恶意行为的新型后门攻击。与依赖训练数据投毒或模型训练操作的传统后门攻击不同，QuRA仅通过量化操作工作。具体而言，QuRA首先采用一种新颖的权重选择策略来识别影响后门目标的关键权重（同时考虑保持模型整体性能）。然后，通过优化这些权重的舍入方向，我们在不降低准确率的情况下跨模型层放大后门效应。大量实验表明，QuRA在大多数情况下实现了接近100%的攻击成功率，且性能下降可忽略不计。此外，我们证明QuRA能够适应并绕过现有的后门防御措施，凸显了其威胁潜力。我们的研究结果强调了广泛使用的模型量化过程中的关键漏洞，强调了需要更强大的安全措施。我们的实现可在</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/cxx122/QuRA" target="_blank">https://github.com/cxx122/QuRA</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">获取。</span></span></p><p cid="n822" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s113-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s113-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n824" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">199、RoundRole: Unlocking the Efficiency of Multi-party Computation with Bandwidth-aware Execution</span></span></p><p cid="n825" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在隐私保护分布式计算系统如安全多方计算(MPC)中，跨方通信是主要瓶颈。过去二十年间，众多卓越协议被提出以降低整体通信复杂度，显著缩小了MPC与明文计算之间的差距。然而，这些进展常常忽视了一个关键方面：</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">非对称</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">通信模式。这种不平衡导致执行过程中产生大量带宽浪费，从而&#34;锁定&#34;了性能。本文提出了RoundRole，一种针对秘密共享MPC的带宽感知执行优化。其核心思想是将决定通信模式的逻辑角色与决定带宽的物理节点</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">解耦</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">。通过将整体协议划分为并行任务，并为每个任务将每个逻辑角色战略性地映射到物理节点，RoundRole能够根据固有协议通信量和物理带宽有效分配通信工作负载。这种执行级别的优化充分利用了网络资源并&#34;解锁&#34;了效率。我们将RoundRole集成到广泛使用的开源MPC框架ABY3之上。在六种不同网络环境（具有同构和异构带宽）下对九种协议进行的广泛评估展示了显著的性能提升，最高可达7.1倍的加速比。</span></span></p><p cid="n826" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f52-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f52-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n828" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">200、RTCON: Context-Adaptive Function-Level Fuzzing for RTOS Kernels</span></span></p><p cid="n829" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">实时操作系统(RTOS)因其包含蓝牙和Wi-Fi等多种子系统而被广泛应用于嵌入式系统。随着其功能不断增长，其攻击面也随之扩大，使其面临更多的安全威胁。为应对这一问题，模糊测试等动态测试技术已被广泛应用于嵌入式系统。然而，对于RTOS，由于其复杂性，这些技术难以有效测试内核中深度嵌套的函数。在本文中，我们提出了RTCon，一种面向RTOS内核的上下文自适应函数级模糊测试工具。RTCon通过在模糊测试过程中自适应生成函数上下文，对RTOS内核中的任何目标函数进行函数级模糊测试。此外，RTCon采用多层分类方法根据置信度对崩溃进行分类，帮助分析师专注于高置信度崩溃。我们实现了RTCon的原型，并在四种流行的RTOS内核上进行了评估：Zephyr、RIOT、FreeRTOS和ThreadX。结果表明，RTCon发现了27个漏洞，其中包括25个新漏洞。我们向维护者报告了所有这些漏洞，并获得了14个CVE编号。RTCon在崩溃分类方面也展示了其有效性，高置信度崩溃的精确度达到92.7%，而低置信度崩溃的精确度仅为5.8%。</span></span></p><p cid="n830" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1600-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1600-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=5b7f3597&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486060%26idx%3D2%26sn%3D75c9796f6cfd6cf0ea4c4ffd390dd333">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sun, 01 Mar 2026 14:04:00 +0800</pubDate>
    </item>
    <item>
      <title>NDSS 2026论文清单及摘要（下）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486060&amp;idx=3&amp;sn=b2d8e7d8dbdb8671d797ec81df76296f</link>
      <description></description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span>2026-03-01 14:04</span> <span style="display: inline-block;">广东</span></p>






  
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=cec8f99c&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fsz_mmbiz_jpg%2FtJDT9c8t2swicHOqPapFS5l3o9ibmW1PkQ5YFbep07v6guMYWhUVlbJQtnElT8MS0JQmhIpic1zMGh19SriaV8wpRkibcKoBuXsr1J5qVia3P7eq8%2F0%3Fwx_fmt%3Djpeg"/></p>
  
  <p class="mp_profile_iframe_wrp" nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="0" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><p cid="n832" mdtype="paragraph" style="box-sizing: border-box;" data-pm-slice="0 0 []"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">201、RTrace: Towards Better Visibility of Shared Library Execution</span></span></p><p cid="n833" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">软件供应链安全近年来已成为一个关键问题。现代软件系统越来越多地依赖第三方依赖项来加速开发。共享库是现代软件系统中软件共享的主要形式，因此也是第三方依赖的主要形式。随着越来越多的攻击针对软件供应链，理解这些依赖项的行为对于识别漏洞和恶意代码至关重要。因此，准确追踪共享库内的函数调用对于有效的软件安全分析至关重要。然而，现有的库函数追踪工具往往无法满足这一需求。正如我们在本文中所展示的，最先进的库函数追踪工具在有效性和可扩展性方面存在局限，遗漏了大量函数调用，并且在处理更复杂的工作负载时失败，导致对运行时行为的不完整或误导性视图。在本文中，我们提出了RTrace，一个旨在解决现有解决方案局限性的追踪工具。我们分析了广泛使用的追踪工具遗漏函数调用的根本原因，并确定了常见陷阱，如依赖不正确的符号信息以及无法监控早期或间接的函数调用。RTrace通过结合全面的运行时监控、函数边界检测以及对隐式和非传统函数调用的支持，克服了这些挑战。我们将RTrace与四种最先进的追踪工具（即ltrace、drltrace、ldaudit和IntelPT）进行了比较。我们在21个应用程序和92个共享库上的评估表明，RTrace在检测函数调用方面显著优于现有工具。RTrace在所有基准测试中至少达到0.92的F1分数，而最好的现有追踪工具仅达到0.74，从而提供了对共享库运行时行为的更准确可见性。最后，我们展示了如何通过提供更完整的共享库函数使用视图，利用RTrace辅助检测恶意包和进行漏洞分析。</span></span></p><p cid="n834" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1243-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1243-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n836" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">202、SACK: Systematic Generation of Function Substitution Attacks Against Control-Flow Integrity</span></span></p><p cid="n837" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">控制流完整性（CFI）是一种广泛采用的防御控制流劫持攻击的技术，旨在限制间接控制传输到一组合法目标。然而，即使在精确的静态CFI策略下，攻击者仍然可以通过函数替换攻击（Sub攻击）来劫持控制流，即用一个仍然允许集合内的有效目标替换另一个有效目标。尽管先前的研究已经通过手动构建证明了此类攻击的可行性，但没有一种方法能够系统化、可扩展地端到端地构建这些攻击。在这项工作中，我们提出了SACK，这是第一个用于大规模自动构建Sub攻击的系统框架。SACK从良性执行中收集触发的间接调用目标，并在大型语言模型的协助下合成安全预言机。然后，它自动执行目标替换，并利用安全预言机检测安全违规，同时确保执行严格遵循精确的CFI策略。我们将SACK应用于七种广泛使用的应用程序，成功构建了419个危及关键安全功能的Sub攻击。我们进一步基于SQLite3、V8和Nginx中的历史漏洞开发了五个端到端漏洞利用程序，实现了任意命令执行或身份验证绕过。我们的研究结果表明，SACK提供了一个可扩展且自动化的管道，能够在不同应用程序中发现大量端到端攻击。</span></span></p><p cid="n838" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2317-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2317-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n840" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">203、SAGA: A Security Architecture for Governing AI Agentic Systems</span></span></p><p cid="n841" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于大型语言模型(LLM)的代理日益自主地相互交互、协作和委派任务，且与人类的互动最小化。代理系统治理的行业指南强调用户需要对其代理保持全面控制，以减轻恶意代理可能造成的潜在损害。多项提出的代理系统设计解决了代理身份、授权和委派问题，但仍停留在纯理论层面，缺乏具体实现和评估。最重要的是，它们不提供用户控制的代理管理。为解决这一差距，我们提出了SAGA，即可扩展的代理系统安全治理架构，使用户能够监督其代理的整个生命周期。在我们的设计中，用户向中央实体(提供者)注册其代理，该实体维护代理的联系信息、用户定义的访问控制策略，并帮助代理在代理间通信中执行这些策略。我们引入了一种用于派生访问控制令牌的密码学机制，可对代理与其他代理的交互进行细粒度控制，并提供正式的安全保证。我们在多个代理任务上评估了SAGA，使用位于不同地理位置的代理以及多种设备端和云端LLM，结果表明在广泛条件下，系统性能开销最小，且不影响底层任务效用。我们的架构实现了自主代理的安全可信部署，促进了该技术在敏感环境中的负责任应用。</span></span></p><p cid="n842" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s869-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s869-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n844" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">204、Scalable Off-chain Auction</span></span></p><p cid="n845" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">区块链拍卖在数字资产（如NFT）的价格发现中发挥着重要作用。然而，尽管其重要性显著，但在以太坊等区块链上直接实施拍卖会面临可扩展性问题。具体而言，链上交易的数量随竞标者数量增加而急剧下降，导致网络拥堵、交易费用上升和交易确认时间延长。这种可扩展性的缺失严重限制了系统处理当今经济中常见的大规模、高速拍卖的能力。在本工作中，我们构建了一个协议，使得拍卖商可以完全在链下进行密封投标拍卖，当各方行为诚实时；如果在n方拍卖协议中有k个竞标者偏离（例如，不公开其密封投标），则链上复杂度仅为O(k)。这优于现有解决方案，即使只有一个竞标者偏离协议，现有解决方案也需要O(n)的链上复杂度。在拍卖商恶意的情况下，我们的协议仍能确保拍卖成功终止。我们实现了该协议，并证明与现有链上解决方案相比，它提供了显著的效率提升。我们使用零知识简洁非交互知识论证（zkSnark）来实现可扩展性，这也确保了链上合约和其他参与者无法获取竞标者身份及其各自投标的信息，除了获胜者和中标金额。</span></span></p><p cid="n846" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s410-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s410-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n848" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">205、SECV: Securing Connected Vehicles with Hardware Trust Anchors</span></span></p><p cid="n849" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代车辆将车外网络（EVN）与车内网络（IVN）集成，以支持导航、诊断和空中更新。这种融合引入了EVN平台作为IVN网关处控制消息的新来源，打破了网关仅过滤来自简单、孤立且隐式信任的遗留ECU流量的传统假设。相反，EVN平台托管了一个具有完整操作系统和多个应用程序的复杂EVN管理器，大大扩大了攻击面：被攻破的操作系统或应用程序可以伪造规避网关过滤的控制消息。我们提出了SECV，一种运行时安全机制，使IVN网关能够准确验证源自EVN的控制消息，即使EVN管理器被攻破。sys在可信执行环境（TEE）内调解所有EVN到IVN的流量，执行每应用程序验证，并附加密码学证明。这些证明由IVN网关使用硬件安全模块（HSM）进行验证，提供低开销的可靠消息认证。SECV解决了TEE-HSM信任建立、实时调解和妥协情况下的稳健归属等实际挑战。在配备ARM TrustZone和符合EVIT标准的HSM的汽车SoC上实现，SECV仅提供6.5%的传输几何平均开销和极端通信突发期间的1.5%额外消息丢失，强制执行强大的安全保证，有效缓解源自EVN的攻击，同时满足实时约束。</span></span></p><p cid="n850" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f106-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f106-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n852" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">206、Select-Then-Compute: Encrypted Label Selection and Analytics over Distributed Datasets using FHE</span></span></p><p cid="n853" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">私有集合交集（PSI）协议允许查询者确定数据集中是否存在某项，而无需揭示查询内容或暴露不匹配的记录。它在欺诈检测、合规监控、健康分析和跨分布式数据源的安全协作等领域有广泛应用。在这些情况下，通过PSI获得的结果可能是敏感的，甚至在向查询者揭示结果之前，需要对相关数据进行某种形式的下游计算，这些计算可能涉及浮点运算，例如机器学习模型的推理。尽管已经提出了许多此类协议，其中一些甚至支持在分布式加密集合上进行安全查询，但它们未能解决上述现实世界的复杂问题。在这项工作中，我们首次提出了&#34;加密标签选择和分析&#34;协议构建，它允许查询者安全地检索不仅限于标识符之间的交集结果，还包括与相交标识符相关联的数据/标签的下游函数结果。为此，我们构建了一种基于近似CKKS全同态加密的新颖协议，支持对实值数据进行高效的标签检索和下游计算。此外，我们引入了几种技术来处理大域中的标识符（例如64位或128位），同时确保下游计算的高精度。最后，我们实现了并基准测试了我们的协议，将其与最先进的方法进行比较，并在真实世界的欺诈数据集上进行了评估，展示了其在大规模用例场景中的可扩展性和效率。我们的结果显示比先前方法快1.4倍至6.8倍，能够在65秒内对真实数据集上的加密标签进行选择和分析，使我们的协议在实际部署中具有实用性。</span></span></p><p cid="n854" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f207-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f207-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n856" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">207、Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference</span></span></p><p cid="n857" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">键值（KV）缓存通过存储中间注意力计算（键值对）以避免冗余计算，是加速大语言模型（LLM）推理的基本机制。然而，这种效率优化引入了显著但尚未充分探索的隐私风险。本文首次对这些漏洞进行了全面分析，证明攻击者可以直接从KV缓存中重建敏感的用户输入。我们设计并实现了三种不同的攻击向量：直接逆向攻击、适用范围更广且更强大的碰撞攻击，以及基于语义的注入攻击。这些方法证明了KV缓存隐私泄露问题的实际性和严重性。为缓解这一问题，我们提出了KV-Cloak，一种新颖、轻量且高效的防御机制。KV-Cloak使用基于可逆矩阵的混淆方案，结合算子融合，来保护KV缓存。我们的广泛实验表明，KV-Cloak有效阻止了所有提出的攻击，将重建质量降低到随机噪声水平。重要的是，它在几乎不损害模型准确性的情况下实现了这种强大的安全性，且性能开销极小，为可信LLM部署提供了实用的解决方案。</span></span></p><p cid="n858" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f258-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f258-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n860" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">208、Should I Trust You? Rethinking the Principle of Zone-Based Isolation DNS Bailiwick Checking</span></span></p><p cid="n861" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DNS缓存投毒攻击通过向解析器中注入伪造的资源记录来秘密劫持域名访问。为应对此类攻击，解析器采用辖区检查（bailiwick checking）这一关键防御机制，旨在过滤DNS响应中可能存在的恶意记录。然而，在第三方服务背景下，域名所有权与传统自上而下的区域授权模型之间的不匹配，对辖区检查的有效性构成了重大挑战。本文对辖区检查的设计与实现进行了系统性分析，证明主流解析器普遍采用保守原则：它们会缓存任何满足最低约束的资源记录，而不管其与原始查询的直接相关性如何。基于这一发现，我们提出了一种新型缓存投毒攻击（称为&#34;布谷鸟域名&#34;）：攻击者通过控制单个子域名，可危害其父域名或兄弟域名。测试结果表明，包括BIND9和Microsoft DNS在内的七种主要DNS解析器实现存在漏洞。通过大规模测量研究，我们确认44.64%的开放解析器和21家主要公共DNS服务提供商也面临风险。此外，我们发现No-IP、ClouDNS和Akamai等7家提供商提供的超过百万个子域名可能易受此类攻击劫持。我们已进行了负责任的披露，向受影响的软件供应商和服务提供商报告了相关问题。BIND9、Unbound、PowerDNS和Technitium已确认我们的报告并分配了3个CVE编号。我们呼吁社区和软件厂商共同应对现代服务生态系统对辖区检查有效性提出的新挑战。</span></span></p><p cid="n862" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f330-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f330-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n864" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">209、Side-channel Inference of User Activities in AR/VR Using GPU Profiling</span></span></p><p cid="n865" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">过去十年，AR/VR设备彻底改变了我们与数字世界的交互方式。用户经常在这些设备上安装的第三方应用中分享敏感信息，如位置、浏览历史，甚至是财务数据，并假设这些信息受到恶意行为者的保护，处于安全环境中。最近的研究表明，恶意应用可以利用这些功能监控良性应用，通过跟踪用户活动，利用性能计数器API等细粒度分析工具。然而，并非所有AR/VR设备（如Meta Quest）都支持应用间监控，因为它们禁用了并发独立应用执行。在本文中，我们提出了OVRWatcher，一种面向AR/VR设备的新型侧通道原语，它通过后台脚本监控低分辨率（1Hz）的GPU使用情况来推断用户活动，这与依赖高分辨率分析的前期工作不同。OVRWatcher能够捕捉不同速度、距离和渲染场景下GPU指标与3D对象交互之间的相关性，无需并发应用执行、应用数据访问或额外SDK安装。我们证明了OVRWatcher在识别独立AR/VR和WebXR应用方面的有效性。OVRWatcher还能区分虚拟对象，例如沉浸式购物应用中真实用户选择的产品以及虚拟会议的参与者数量，从而揭示用户的产品偏好并可能暴露会议中的机密信息。OVRWatcher在应用识别方面实现了超过99%的准确率，在对象级推断方面实现了超过98%的准确率。</span></span></p><p cid="n866" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1302-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1302-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n868" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">210、SIPConfusion: Exploiting SIP Semantic Ambiguities for Caller ID and SMS Spoofing</span></span></p><p cid="n869" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">会话初始化协议（SIP）是现代实时通信系统的基石，为VoIP、VoLTE和RCS等服务中的语音通话、文本消息和多媒体会话提供支持。尽管SIP提供了身份验证和身份声明的机制，但其固有的灵活性可能导致不同实现之间存在语义歧义，从而被攻击者利用。在本文中，我们提出了SIPChimera，一种新颖的黑盒模糊测试框架，旨在系统性地识别SIP实现中基于身份歧义的身份欺骗漏洞。我们对六种广泛使用的开源SIP服务器（包括Asterisk和OpenSIPS）和九种流行的用户代理进行了SIPChimera评估，发现攻击者可以通过操纵身份头信息来欺骗身份并绕过身份验证。我们通过评估五种VoIP设备、七种商业SIP部署和三种运营商级基于RCS的短信平台，展示了这些漏洞的现实影响。我们的实验表明，攻击者可以利用这些漏洞在VoIP通话中进行来电显示欺骗，并通过RCS发送欺骗性短信，冒充任意用户或服务。我们已向相关供应商负责任地披露了我们的发现，并收到了积极确认。最后，我们提出了缓解这些问题的解决方案。</span></span></p><p cid="n870" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s116-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s116-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n872" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">211、Small Cell, Big Risk: A Security Assessment of 4G LTE Femtocells in the Wild</span></span></p><p cid="n873" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">femtocell是小型、由运营商部署的基站，旨在扩展移动网络覆盖范围，但它们与运营商移动基础设施的整合引入了显著的新攻击面。虽然5G femtocell标准最近才最终确定，但4G LTE femtocell已经标准化并广泛实施。在这项工作中，我们基于真实商业设备和大规模互联网测量，对4G LTE femtocell进行了首次系统性安全评估。我们系统分析了4款商业femtocell设备的软件和硬件，确定了5个关键且普遍存在的漏洞，这些漏洞可导致本地或远程系统被攻破。我们的全球互联网测量发现了86,108个疑似femtocell部署，其中许多容易受到远程攻击。此外，我们在真实运营商网络中实验验证了单个被攻破的femtocell可作为攻击移动核心网络及其订阅者的有力入口点。我们的研究结果表明，在现有4G LTE网络中，femtocell安全仍然是一个紧迫的关切问题。我们将研究结果报告给了全球移动通信系统协会（GSMA）和第三代合作伙伴计划（3GPP）服务与系统方面工作组3（SA3）。3GPP SA3随后批准了一项进一步强化5G femtocell安全的研究项目，以及一项定义5G femtocell安全保证规范（SCAS）的工作项目。</span></span></p><p cid="n874" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1968-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1968-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n876" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">212、SNPeek: Side-Channel Analysis for Privacy Applications on Confidential VMs</span></span></p><p cid="n878" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">基于可信执行环境（TEEs）的机密虚拟机（CVMs）能够实现新的隐私保护解决方案。然而，它们将侧信道泄漏排除在其威胁模型之外，将缓解此类攻击的责任转移给开发者。但是，这些缓解措施要么不够通用，要么在实际应用中速度太慢，而且开发者目前缺乏一种系统、高效的方法来测量和比较实际部署中的泄漏情况。在本文中，我们提出了SNPeek，一个开源工具包，它可在生产级AMD SEV-SNP硬件上提供可配置的侧信道跟踪原语，并结合基于统计和机器学习的分析流程，实现自动化的泄漏估计。我们将SNPeek应用于三个部署在CVM上以增强用户隐私的代表性工作负载——私有信息检索、私有频繁项和Wasm用户定义函数，并发现了先前未被注意到的泄漏，包括一个以497 kbit/s速率泄露数据的隐蔽信道。结果表明，SNPeek能够精确定位漏洞，并指导基于 oblivious memory 和差分隐私的低开销缓解措施，为从业者提供了一条具有实际意义的部署具有实质性保密保证的CVMs的路径。</span></span></p><p cid="n879" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f699-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f699-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n881" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">213、SoK: Analysis of Accelerator TEE Designs</span></span></p><p cid="n882" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">加速器可信执行环境（TEE）是一种流行技术，为加速器中的敏感数据/代码提供强大的机密性、完整性和隔离保护。然而，大多数研究针对特定CPU或加速器设计，因此缺乏通用性。最近的TEE调查部分总结了加速器计算中的威胁和保护措施，但尚未提供构建加速器TEE的指南，也未比较其安全解决方案的优缺点。本文多年来对加速器TEE进行了全面分析。我们总结了构建加速器TEE的典型框架，并归纳了从软件到物理攻击的广泛使用的攻击向量。此外，我们对加速器TEE的三大安全机制进行了系统化：(1)访问控制，(2)内存加密/解密，(3)认证。对于每个方面，我们比较了现有研究中不同的安全解决方案并总结了它们的见解。最后，我们分析了影响TEE在实际平台部署的因素，特别是关于可信计算基（TCB）和兼容性问题。</span></span></p><p cid="n883" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1424-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1424-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n885" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">214、SoK: Cryptographic Authenticated Dictionaries</span></span></p><p cid="n886" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">我们对认证字典（ADs）的研究进行了系统化整理——这是一种密码学数据结构，能够支持密钥透明度、二进制透明度、可验证键值存储以及完整性保留文件系统等应用。首先，我们提出了一个统一框架，概括了五种常见部署场景背后的信任和威胁假设。其次，我们提炼并调和了文献中分散的各种安全定义，明确了它们提供的保证以及各自的适用场景。第三，我们构建了AD结构的分类法，并分析了它们的渐近成本，揭示了一个明显的二元对立：所有已知方案要么在查找和更新操作上都需O(log n)时间，要么仅通过使另一操作付出O(n)的代价来实现某一操作的O(1)时间复杂度。令人惊讶的是，即使引入更强的信任假设，这一障碍仍然存在，这削弱了&#34;更多信任换取效率&#34;的直观认识。最后，我们提出了应用驱动的研究问题，包括现实的审计模型以及在当前完全不提供可验证完整性的系统中促进采用的激励机制。</span></span></p><p cid="n887" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1465-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1465-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n889" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">215、SoK: Take a Deep Step into Linux Kernel Hardening Effectiveness from the Offensive-Defensive Perspective</span></span></p><p cid="n890" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">尽管人们已付出巨大努力来加固Linux内核——这一支撑众多广泛使用的发行版（如Ubuntu、Debian、Fedora）的基础——但它仍然持续面临复杂且顽固的内存安全漏洞。在本研究中，我们引入了一个新颖的系统性框架，从攻击者的角度将内核利用分解为三个不同阶段。通过对2015年以来121个公开记录的漏洞利用进行综合分析，我们识别并分类了64个反复出现的攻击向量。利用这种结构化方法，我们对51个现有的内核防御机制进行了深入评估，清晰地映射了它们的覆盖范围、局限性、冗余性和相互依赖性。我们的研究结果揭示了显著的保护缺口：23个攻击向量完全没有得到保护，31个现有防御机制可以被绕过或已过时。此外，我们还发现流行下游发行版在理论有效性与实际部署之间存在显著差异，突显了四个主要发行版中4个未被充分利用的加固措施和配置错误。通过阐明这些关键缺口并提供可行的见解，我们的工作指导内核开发者和安全实践者加强防御策略并完善未来安全设计。</span></span></p><p cid="n891" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1725-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1725-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n893" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">216、SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities</span></span></p><p cid="n894" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">传感器是信息物理系统(CPS)的基础，通过将物理刺激转换为数字测量值，实现感知和控制。然而，尽管对传感器物理攻击的研究日益增多，由于该领域的临时性特点，我们对传感器硬件漏洞的理解仍然零散。此外，无限的攻击信号空间进一步威胁抽象和防御复杂化。为解决这一差距，我们提出了一个系统化框架，称为传感器带外(OOB)漏洞，首次基于底层物理原理为传感器攻击面提供了全面抽象。我们采用自底向上的系统化方法，分析三个层面的OOB漏洞。在组件层面，我们确定导致OOB漏洞的物理原理和局限性。在传感器层面，我们对已知攻击进行分类并评估其实用性。在系统层面，我们分析传感器融合、闭环控制和智能感知等CPS特性如何影响OOB威胁的暴露和缓解。我们的研究结果为传感器硬件安全提供了基础理解，并为旨在构建更安全传感器和CPS的传感器设计师、安全研究人员和系统开发者提供了指导及未来方向。</span></span></p><p cid="n895" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s450-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s450-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n897" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">217、STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production</span></span></p><p cid="n898" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">模型参数和用户数据的隐私对于基于Transformer的云服务（如在线聊天机器人）至关重要。虽然最近在安全多方计算和同态加密方面的进展提供了强大的密码学保证，但其计算开销使得它们对于大规模Transformer模型的实时推理变得不可行。在这项工作中，我们提出了一种实用的替代方案，在实际部署中平衡隐私和效率。我们引入了一个三方威胁模型，涉及模型开发者、云模型服务器和数据所有者，捕捉了实际AI服务的信任假设和部署条件。在该框架内，我们设计了一种基于半对称置换的保护机制，并提出了STIP，这是首个可在商用硬件上部署的大规模Transformer三方隐私保护推理系统。STIP在保持无损推理准确性的同时，正式限制了隐私泄露。为进一步保护模型参数，STIP集成了可信执行环境以抵御模型提取和微调攻击。我们在六种代表性的Transformer模型家族（包括多达700亿参数的模型）和三种部署设置下评估了STIP。STIP的效率与无保护的全云推理相当，例如，STIP在LLaMA2-7B模型上实现了31.7毫秒的延迟。STIP还表现出对用户数据和模型参数各种攻击的有效抵抗力。STIP已在我们专有的70B模型的生产环境中部署。在为期三个月的在线测试中，STIP仅带来12%的额外延迟，且未报告任何隐私事件，证明了其在生产规模AI系统中的实用性和稳健性。</span></span></p><p cid="n899" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s35-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s35-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n901" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">218、Strategic Games and Zero-Shot Attacks on Heavy-Hitter Network Flow Monitoring</span></span></p><p cid="n902" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">高频检测是线级DDoS缓解和速率限制的基础，然而其针对自适应攻击者的鲁棒性在很大程度上尚未被探索。我们构建了一个端到端评估框架，将高频检测逻辑嵌入到交换级模拟器中，并使用强化学习自动调整其参数，以对网络中的大象流进行速率限制。随后，我们将该保护系统与一个自适应攻击者对抗，该攻击者学习在规避检测的同时最大化吞吐量，并展示其能够将配置的带宽上限提高高达299%，暴露了系统性的盲点。为了加强监控系统，我们采用了一种联合对抗训练形式：检测器与攻击者共同进化，达到一种攻防纳什均衡，其中攻击者利用网络带宽的能力降低了2.2倍。最后，我们证明可以使用机器学习创建智能数据包合成器，能够在9个测试系统中的8个上执行带宽利用，而无需针对检测系统的任何先验知识。我们将其称为零次攻击，因为它不需要了解目标高频检测系统即可执行其功能。我们的开源框架有助于量化未被充分照亮的攻击面，并为对抗鲁棒的数据平面流监控提供了一种建设性方法。</span></span></p><p cid="n903" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1301-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1301-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n905" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">219、Success Rates Doubled with Only One Character: Mask Password Guessing</span></span></p><p cid="n906" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">虽然传统的全密码猜测攻击已被广泛研究，但很少有研究探索掩码密码猜测，即攻击者通过利用各种侧信道攻击（如肩窥、指纹和按键音频反馈）以某种方式获得了目标受害者的密码的部分信息（如长度和/或某些字符）。为了评估具有不同能力的掩码攻击者构成的威胁，我们研究了四种主要的掩码猜测场景，每种场景基于攻击者利用的不同类型的信息（例如，受害者密码的长度和某些字符）。我们首次通过提出两种密码模型（基于神经网络的PassSeq和基于概率统计的Kneser-Ney），系统地全面地描述了结合侧信道先验、可识别个人信息（PII）和先前泄露的（姐妹）密码的掩码猜测的影响。我们使用最大似然估计技术，提出了一种新的猜测次数估计方法，以准确高效地估计在给定密码模型下针对目标密码所需的猜测次数。在15个大规模数据集上的广泛实验证明了PassSeq和Kneser-Ney的有效性。特别是在十次猜测内：（1）当拖网攻击者知道受害者4位PIN码的字符组成（无顺序）时，成功率提高152%（从14%增至35%）；（2）当基于PII的定向攻击者知道受害者密码的长度时，成功率提高47%-82%；（3）如果该定向攻击者还知道受害者密码的一个字符（除长度外），成功率通常翻倍，达到7%-29%（而对于能够利用受害者姐妹密码的定向攻击者，这些数字将达到33%-73%）。为了进一步验证我们掩码猜测模型的实用性，我们从11种流行键盘（如苹果、戴尔、联想）收集了真实的按键音频数据，并复制了通过声学侧信道推断部分密码信息的攻击。实验表明，我们的PassSeq显著提高了现有按键推断攻击的成功率，在10次猜测内实现了额外5.6%-166.7%的改进。这项工作强调掩码密码猜测是一种值得更多关注的破坏性威胁。</span></span></p><p cid="n907" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1059-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1059-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n909" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">220、SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition</span></span></p><p cid="n910" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">联邦学习（FL）能够在不共享原始数据的情况下实现协作模型训练，但容易受到梯度反转攻击（GIA），即攻击者通过共享的梯度重建私有数据。现有防御方法要么对嵌入式平台造成不切实际的计算开销，要么无法同时实现隐私保护和良好的模型效用。此外，许多防御方法可以被已获取防御细节的自适应攻击者轻易绕过。为解决这些局限性，我们提出了SVDefense，一种针对GIA的新型防御框架，利用截断奇异值分解（SVD）来模糊梯度更新。SVDefense引入了三项关键创新：自适应能量阈值，能够适应客户端的脆弱性；通道加权近似，选择性保留有效模型训练所需的关键梯度信息，同时增强隐私保护；以及层加权聚合，用于处理类别不平衡情况下的有效模型聚合。我们的广泛评估表明，在图像分类、人体活动识别和关键词识别等多个应用中，SVDefense通过提供强大的隐私保护且对模型精度影响最小，优于现有防御方法。此外，SVDefense适用于部署在各种资源受限的嵌入式平台上。论文接受后，我们将公开我们的代码。</span></span></p><p cid="n911" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s114-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s114-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n913" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">221、SYSYPHUZZ: the Pressure of More Coverage</span></span></p><p cid="n915" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">内核模糊测试能有效发现漏洞。虽然现有的内核模糊测试器主要专注于最大化代码覆盖率，但仅靠覆盖率并不能保证彻底的探索。此外，旨在最大化覆盖率的现有模糊测试器已进入平台期。这一紧迫情况凸显了需要一个新的方向：面向代码频率的内核模糊测试。然而，增加对低频内核代码的探索面临两个关键挑战：(1)资源限制使得在不导致任务爆炸的情况下难以调度足够的任务来探索低频区域。(2)随机突变常常会破坏针对低频区域的系统调用上下文依赖，降低模糊测试的有效性。在我们的论文中，我们首先通过评估Syzkaller在Linux内核中的表现，对不平衡代码覆盖率进行了细粒度研究，并作为回应，提出了SYSYPHUZZ，一个旨在增强对测试不足代码区域探索的内核模糊测试器。SYSYPHUZZ引入了选择性任务调度，以动态优先排序和管理探索任务，避免任务爆炸。它还采用上下文保持突变策略，降低破坏重要执行上下文的风险。我们将SYSYPHUZZ与最先进的(SOTA)内核模糊测试器Syzkaller和SyzGPT进行了比较评估。我们的结果表明，SYSYPHUZZ显著减少了探索不足的代码区域数量，发现了Syzkaller遗漏的31个独特漏洞和SyzGPT遗漏的27个漏洞。此外，SYSYPHUZZ还发现了Syzbot遗漏的5个漏洞，Syzbot在数百台虚拟机上持续运行，这证明了SYSYPHUZZ的有效性。为了评估SYSYPHUZZ对最先进模糊测试器的增强效果，我们将它与SyzGPT集成，产生了SyzGPTsysy，它发现了多33%的独有漏洞，凸显了SYSYPHUZZ的潜力。所有发现的漏洞都已负责任地披露给Linux维护者。我们在</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/HexHive/Sysyphuzz" target="_blank">https://github.com/HexHive/Sysyphuzz</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">上发布了SYSYPHUZZ的源代码，并正在尝试将其合并到Syzkaller中。</span></span></p><p cid="n916" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s921-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s921-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n918" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">222、Targeted Password Guessing Using k-Nearest Neighbors</span></span></p><p cid="n919" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着用户密码账户数量的不断增加，用户越来越倾向于重复使用密码。最近，已有大量研究致力于构建针对性的密码猜测模型来表征用户的密码重复使用行为。然而，现有研究主要专注于通过仅训练相似的密码对（例如，textnormal{texttt{Shark0301} → texttt{shark03}}）来表征微小的修改行为。这导致了过拟合问题，使得现有模型忽视了用户的大幅度修改行为（例如，textnormal{texttt{Shark0301} → texttt{Bear03}}）。为填补这一空白，本文引入了一种名为 emph{k}-最近邻针对性密码猜测（KNN-TPG）的新非参数方法。KNN-TPG构建了一个数据存储，保留了所有源密码的上下文向量以及目标密码的前缀。在生成新密码的过程中，KNN-TPG从数据存储中检索 emph{k} 个最近邻向量，以确保生成的密码更好地符合真实的密码分布。通过创造性地将KNN-TPG与我们提出的基于Transformer的密码模型相结合，我们提出了一个新的针对性密码猜测模型，即KNNGuess。在生成新密码的每一步，KNNGuess预测并利用三种不同的分布，旨在全面建模用户的密码重复使用行为。我们通过大量实验验证了KNNGuess模型和KNN-TPG方法的有效性，这些实验包括12个大规模真实世界密码数据集，包含48亿个密码。更具体地说，当用户在网站A的密码（即$pw_A$）被泄露时，在100次猜测内，KNNGuess猜测其在网站B的密码（即$pw_B$，且$pw_B$$neq$$pw_A$）的成功率对于普通用户为25.40%，对于安全意识较强的用户为10.26%，比其主要竞争对手高出8.52%-119.0%（平均55.33%）。与最先进的密码模型（即Pass2Edit和PointerGuess）相比，这一数值高出8.52%-27.66%（平均18.09%）。我们的研究结果表明，密码篡改攻击的威胁比用户预期的要高。</span></span></p><p cid="n920" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s2077-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s2077-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n922" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">223、Targeted Physical Evasion Attacks in the Near-Infrared Domain</span></span></p><p cid="n923" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">多种攻击依赖于红外光源或吸热材料，以在各种图像识别应用中不可察觉地欺骗系统，使其错误解读视觉输入。然而，几乎所有现有方法只能发起无目标攻击，并且由于用例特定约束（如位置和形状）而需要大量优化。本文提出了一种新颖、隐蔽且经济高效的攻击方法，能够生成有目标和无目标的对抗性红外扰动。通过使用现成的红外手电筒将透明薄膜上的投影投射到目标物体上，我们的方法首次能够在红外领域可靠地发起无激光有目标攻击。在数字和物理领域交通标志上的大量实验表明，与先前工作相比，我们的方法在各种攻击场景中（包括不同光照条件、距离和角度）具有更强的鲁棒性，并能取得更高的攻击成功率。同样重要的是，我们的攻击方法成本极低，部署成本不到50美元，仅需几十秒。最后，我们提出了一种基于分割的新型检测方法，能够有效抵御我们的攻击，F1分数高达99%。</span></span></p><p cid="n924" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1568-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1568-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n926" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">224、TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious Campaigns on X</span></span></p><p cid="n927" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在线社交网络（OSNs）中的恶意行为者使用脚本控制的社会机器人，通过回复或评论与用户互动。这些机器人被编程为仅在帖子中出现特定触发关键词时才激活。我们将这类先进的上下文感知活动者称为触发机器人（TB）代理，其目的是欺骗用户为非法产品付款或泄露敏感的金融凭证。本文对TB代理的检测和特征进行了系统性和数据驱动的研究。我们介绍了TBTrackerX，这是一个为收集和分析TB活动而设计的新框架。使用该系统，我们从2,647个独特的TB代理中捕获了4,452个TB代理回复，这些回复针对我们的蜜罐账户，并揭示了与X平台上超过84K用户的互动。我们的研究结果表明，TB代理通过使用上下文相似的回复（相似度高达0.97）、表现出间歇性发布模式（爆发时间从15秒到5分钟不等）以及在活动高峰期后采用休眠行为来规避检测。此外，我们还识别出一个协调的TB生态系统，其特征是虚假的TB关注者和共享的TB主控者。这项研究强调了迫切需要更好的审核和检测机制来对抗这些复杂的社会媒体操纵形式。</span></span></p><p cid="n928" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1239-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1239-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n930" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">225、The Dark Side of Flexibility: Detecting Risky Permission Chaining Attacks in Serverless Applications</span></span></p><p cid="n931" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">现代无服务器平台通过将基础设施与函数级开发解耦，实现了应用的快速演进。然而，这种灵活性导致无服务器应用的去中心化函数级权限配置与集中式云访问控制系统之间存在根本性不匹配。我们观察到，这种不匹配通常会导致无服务器应用中的函数存在风险权限，攻击者可以利用这些风险权限链接多个函数来提升权限、接管账户，甚至横向移动以入侵其他账户。我们将此类攻击称为&#34;风险权限链接攻击&#34;。在本工作中，我们提出了一种自动化推理系统，能够检测可用于链接攻击的风险权限。首先，我们基于以攻击者为中心的模式抽象方法，明确捕获了来自不同函数和账户的独立权限如何合并为实际的攻击链。基于这种抽象，我们构建了一个模式引导的检测工具，用于发现现实世界无服务器应用中的可利用权限链。我们通过分析来自AWS和阿里云官方生产级应用仓库的无服务器应用，评估了我们的方法。结果表明，我们的分析发现了28个存在漏洞的应用，包括5个已确认的CVE、6个负责任的漏洞认可和1个安全赏金。这些发现表明，风险权限链接攻击不仅是理论风险，也是已经存在于商业无服务器部署中的结构性且可利用的威胁，其根源在于去中心化无服务器应用与集中式访问控制模型之间的根本性不匹配。</span></span></p><p cid="n932" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s819-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s819-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n934" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">226、The Heat is On: Understanding and Mitigating Vulnerabilities of Thermal Image Perception in Autonomous Systems</span></span></p><p cid="n935" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">热成像相机日益被视为自主系统中确保低能见度条件下感知能力的可行解决方案。自动驾驶汽车、机器人和无人机的热感知集成管线中集成了专业光学元件和先进信号处理技术，能够捕捉相对温度变化，并在传统可见光相机难以应对的场景（如夜间、雾天或大雨）中检测生物和物体。然而，热感知系统的安全性和可信度是否与传统相机相当，目前尚不清楚。我们的研究揭示了热图像处理中存在的三种新型漏洞，这些漏洞存在于热相机固有的均衡化、校准和透镜机制中。这些漏洞可由环境中自然存在或恶意放置的热源触发，改变感知到的相对温度，或产生时间控制的人工制品，从而阻碍障碍物避让功能的正常运行。我们系统分析了三种自主系统用热相机（FLIR Boson、InfiRay T2S、FPV XK-C130）中的漏洞，评估了它们对三种微调热物体检测器和两种可见光-热融合自动驾驶模型的影响。研究结果显示，由于均衡化过程中的缺陷，行人检测的平均精度下降了50%，融合模型下降了45%。最高时速40公里的真实道路测试显示，行人误检率高达100%，且能以91%的成功率制造虚假障碍物，这些影响在攻击结束后仍会持续数分钟。为解决这些问题，我们提出了并评估了三种新型威胁感知信号处理算法，能够动态检测并抑制攻击者引入的人工制品。我们的研究结果揭示了热感知过程的可靠性，旨在提高人们对该技术用于障碍物避避时局限性的认识。</span></span></p><p cid="n936" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s330-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s330-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n938" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">227、The Role of Privacy Guarantees in Voluntary Donation of Private Health Data for Altruistic Goals</span></span></p><p cid="n939" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">出于利他目的自愿捐赠私人健康信息，例如支持研究进展，是一种常见做法。然而，对数据滥用和泄露的担忧可能会阻碍人们捐赠其信息。隐私增强技术（PETs）旨在缓解这些担忧，从而实现安全私密的数据共享。本研究通过在Prolific平台招募参与者进行了一项情景调查（N=494），考察了美国人在四种PETs提供的通用保障下，为开发新治疗而捐赠医疗数据的意愿：数据过期、匿名化、目的限制和访问控制。研究探讨了验证这些保障的两种机制：自我审计和专家审计，并控制了混杂因素的影响，包括人口统计特征以及两种类型的数据收集机构：营利性和非营利性机构。我们的研究结果表明，受访者对非营利实体事先抱有极高的隐私期望，因此明确列出隐私保护措施对其整体感知影响甚微。相比之下，提供隐私保障提升了受访者对营利实体的隐私期望，使其与非营利组织的期望几乎持平。此外，尽管技术界建议将审计作为增加对PET保障信任的机制，但我们观察到关于此类审计透明度的效果有限。我们强调了这些发现相关的风险，并强调了未来跨学科研究工作的迫切需要，以弥合技术界与终端用户之间在审计PETs有效性认知方面的差距。</span></span></p><p cid="n940" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s518-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s518-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n942" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">228、There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models</span></span></p><p cid="n943" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）被广泛用于信息获取，但其内容审核行为在不同地理和语言背景下差异显著。本文对来自15个领先LLMs的70多万条回复进行了首次全面分析，这些回复从12个地点使用1,118个涵盖五个类别的敏感查询（涉及13种语言）进行评估。我们发现存在显著的地理差异，审核率在不同地点间相对差异高达60%——例如，软审核（如回避性回复）在德语语境中出现率为14.3%，而在祖鲁语语境中为24.9%。按类别分析，其他（通常不安全）、仇恨言论和性内容比政治或宗教内容受到更严格的审核，其中政治内容显示出最大的地理变异性。我们还观察到在线和离线模型版本之间的差异，例如DeepSeek本地部署时的软审核率比通过API调用时高出15.2%。回复长度（和时间）分析显示，审核过的回复平均比未审核的回复短约50%。这些发现对AI公平性和数字平等具有重要意义，因为不同地点的用户获得的信息访问不一致。我们首次提供了LLM内容审核中地理跨语言偏差的系统证据，并展示了模型选择如何极大影响用户体验。</span></span></p><p cid="n944" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f593-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f593-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n946" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">229、ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking</span></span></p><p cid="n947" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLMs）已成为广泛应用的基础组件，包括自然语言理解与生成、具身智能和科学发现。随着其计算需求的持续增长，这些模型越来越多地被部署为云服务，使用户能够通过互联网访问强大的LLMs。然而，这种部署模式引入了一类新的威胁：通过无限推理的拒绝服务（DoS）攻击，攻击者精心设计输入，导致模型进入过长的或无限生成循环。这些攻击会耗尽后端计算资源，降低或拒绝向合法用户提供服务。为缓解此类风险，许多LLM提供商采用闭源、黑盒设置来隐藏模型内部结构。在本文中，我们提出了ThinkTrap，一种针对LLM服务的DoS攻击的新型输入空间优化框架，即使在黑盒环境中也能实施。ThinkTrap的核心思想是将离散令牌映射到连续嵌入空间，然后在利用输入稀疏性的低维子空间中进行高效的黑盒优化。此优化的目标是识别能够诱导先进LLMs进行延长或非终止生成的对抗性提示，以实现最小的令牌开销的DoS攻击。我们在多个商业闭源LLM服务上评估了所提出的攻击。结果表明，即使在远低于这些平台通常实施的限制性请求频率限制（通常每分钟10次请求，10 RPM）的情况下，该攻击仍可将服务吞吐量降低至原始容量的1%，在某些情况下甚至导致完全服务失效。</span></span></p><p cid="n948" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f639-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f639-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n950" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">230、Through the Authentication Maze: Detecting Authentication Bypass Vulnerabilities in Firmware Binaries</span></span></p><p cid="n951" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">嵌入式Web服务已广泛应用于路由器和网关等网络设备中。这些服务通常暴露在公共网络上，使其成为身份验证绕过攻击的诱人目标。此类漏洞允许攻击者无需有效凭据即可获得特权访问，对设备完整性和网络安全构成严重威胁。现有的检测技术主要依赖手动分析或刚性启发式方法，在面对多样化且不断发展的身份验证方案时效果不佳。我们提出了AuthSpark，一种用于检测固件二进制文件中身份验证绕过漏洞的新型动态分析框架。AuthSpark利用成功和失败身份验证尝试之间的执行轨迹相似性来定位凭据检查点，然后跟踪身份验证相关变量的传播以识别身份验证成功逻辑，最后采用具有特定任务能力调度和变异策略的自定义灰盒模糊测试器来探索绕过路径。我们在32个包含14个已知漏洞的真实设备固件上评估了AuthSpark。AuthSpark成功识别出44个凭据检查中的42个，并检测到所有14个已知漏洞。更重要的是，当应用于最新版本的固件时，AuthSpark发现了6个零日身份验证绕过漏洞，其中4个已获得官方编号（3个CVE和1个PSV）。这些结果凸显了AuthSpark的有效性及其发现真实系统中关键安全漏洞的潜力。</span></span></p><p cid="n952" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2757-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2757-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n954" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">231、Tickets to Hide: An Inside Look into the Anti-Abuse Ecosystem through Internal Abuse Data</span></span></p><p cid="n955" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">各种治理工具旨在打击互联网滥用行为——从删除受版权保护内容的立法到阻止垃圾邮件的屏蔽列表。反过来，这些工具依赖于行业标准来处理滥用行为：向网络所有者报告滥用情况并请求缓解措施。尽管许多托管服务提供商迅速采取行动以保持互联网环境的清洁，但有些则没有。这就引发了一个问题：哪种类型的滥用会得到后续处理，以及决定采取缓解措施或忽略所报告滥用的理由是什么。通过与荷兰执法部门的独特合作，我们获得了进入一家以滥用行为闻名的托管服务提供商运营后端的权限。对其内部滥用处理机制的罕见一瞥使我们能够研究影响反滥用行动的反滥用生态系统中的机制。我们发现，客户通知率高度依赖于报告者和滥用类别。与儿童性虐待材料(CSAM)和垃圾邮件相关的滥用报告会导致采取缓解措施，而关于版权侵权和端口扫描的报告则经常被忽视。诸如屏蔽、解除对等连接和执法部门查询等可能直接影响业务连续性的治理工具会影响客户通知，而个人滥用报告则容易被忽视。我们希望这项研究能够为政策制定者提供参考，使治理工具包与实际的滥用处理实践保持一致。</span></span></p><p cid="n956" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f468-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f468-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n958" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">232、Time and Time Again: Leveraging TCP Timestamps to Improve Remote Timing Attacks</span></span></p><p cid="n960" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">最著名的侧信道攻击之一是通过执行特定操作所需的时间来推断秘密信息。许多系统已被证明容易受到此类攻击，范围从加密算法、Web应用程序到微架构实现。通过网络连接利用这些侧信道泄露已被证明具有挑战性，这是由于往返时间的变化，即网络抖动。随着处理器速度变快导致时间差异变小，系统变得更复杂使得收集一致测量更加困难，以及网络拥塞加剧网络抖动，时序攻击已变得尤其具有挑战性。在这项工作中，我们引入了新的远程时序攻击方法，这些方法完全不受网络路径上的抖动影响，使其比基于往返时间的时序攻击效率提高数倍，并且能够检测到更小的时间差异。更具体地说，执行时间是从服务器在确认请求和发送响应时生成的TCP时间戳值推断出来的。此外，我们展示了如何利用对传入请求的顺序处理来扩展与秘密相关的操作时间，从而实现更准确的攻击。最后，通过广泛的测量和实际案例研究，我们证明了本文介绍的技术与其他时序攻击方法相比具有多种优势：需要更少的前提条件，任何基于TCP的协议都容易受到这些攻击，并且这些攻击可以分布式执行。</span></span></p><p cid="n961" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s893-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s893-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n963" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">233、Time will Tell: Large-scale De-anonymization of Hidden I2P Services via Live Behavior Alignment</span></span></p><p cid="n964" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">I2P（隐形互联网项目）是一种流行的匿名通信网络。尽管现有的I2P去匿名化方法专注于在大量网络流量中识别目标隐藏服务的潜在流量模式，但它们往往无法有效地扩展到由众多路由器组成的大型且多样化的I2P网络。在本文中，我们介绍了一种名为I2PERCEPTION的低成本方法，用于揭示I2P隐藏服务的IP地址。在I2PERCEPTION中，攻击者部署floodfill路由器来被动监控I2P路由器并收集其RouterInfo。我们分析了路由器信息发布机制，以准确识别路由器的加入（即开启）和离开（即关闭）行为，从而实现对I2P网络细粒度的实时行为推断。通过主动探测获取托管在I2P路由器之一上的目标隐藏服务的实时行为（即开启-关闭模式）。通过关联目标隐藏服务和I2P路由器的实时行为，我们缩小了与隐藏服务行为匹配的路由器集合，从而揭示隐藏服务的真实网络身份以实现去匿名化。通过在八个月内仅部署15个floodfill路由器，我们通过大量真实实验验证了我们方法的精确性和有效性。结果表明，I2PERCEPTION成功地去匿名化了所有受控的隐藏服务。</span></span></p><p cid="n965" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f114-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f114-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n967" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">234、TIPSO-GAN: Malicious Network Traffic Detection Using a Novel Optimized Generative Adversarial Network</span></span></p><p cid="n968" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">检测高级网络威胁，特别是零日漏洞，在网络安全中构成重大挑战。本文提出了TIPSO-GAN，一种用于检测恶意流量的优化生成对抗网络（GAN）。TIPSO-GAN通过将GAN训练构建为群体优化问题，利用集体智能进行复杂优化，解决了基于GAN的入侵检测系统（IDS）的常见问题，如训练不稳定性和模式崩溃。为了增强粒子群优化（PSO），TIPSO-GAN采用了三种策略：（1）自适应惯性权重以平衡探索与开发，（2）多样性保持策略以防止过早收敛，（3）反馈循环以重新初始化停滞粒子。TIPSO-GAN将迁移学习与时间衰减多头自注意力机制相结合，以优先考虑近期特征，有助于检测未见过的恶意流量。目标函数中结合重构损失和焦点损失，进一步确保正常样本的真实性，同时关注具有挑战性的恶意样本。在CIC-IDS2018、CICAPT-IIoT2024和CIC-DDoS2019数据集上，TIPSO-GAN分别实现了99.1±0.1、98.9±0.1和98.7±0.1的F1值，比最强基线模型高出0.2-1.0 F1，并超过了transformer IDS模型。在CICAPT-IIoT2024上，它达到了0.999±0.002的宏观PR-AUC，领先于次优方法（0.960±0.005）。在严格的零日评估中，TIPSO-GAN在LOFO测试中达到92.3 F1，在跨数据集实验中达到79-83 F1，同时保持召回率高于0.80。尽管经过PSO增强训练，TIPSO-GAN仍保持0.42毫秒延迟、约2400流/秒吞吐量和2.1 GB内存占用，性能稳定至10^8流。我们的代码可在</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/osampas27/tipsoganmod" target="_blank">https://github.com/osampas27/tipsoganmod</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">获取。</span></span></p><p cid="n969" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f3241-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f3241-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n971" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">235、To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling</span></span></p><p cid="n972" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">差分随机梯度下降（DP-SGD）算法支持以形式化差分隐私（DP）保证训练机器学习（ML）模型。传统上，DP-SGD使用泊松子采样在每个迭代中选择批次来处理训练数据。最近，由于更好的兼容性和更低的计算开销，洗牌已成为一种常见替代方案。然而，在洗牌下计算严格的理论DP保证仍然是一个开放问题。因此，使用洗牌训练的模型通常被评估为好像使用了泊松子采样，这可能导致不正确的隐私保证。这提出了一个引人入胜的研究问题：我们能否验证使用洗牌的最先进模型所报告的理论DP保证与其实际泄漏之间是否存在差距？为此，我们定义了新的DP审计程序来分析使用洗牌的DP-SGD，并衡量它们在不同批次大小、隐私预算和威胁模型下紧密估计隐私泄漏的能力。总体而言，我们证明使用这种方法训练的DP模型大大高估了其隐私保证（高达4倍）。然而，我们也发现理论泊松DP保证与洗牌实际隐私泄漏之间的差距并非在所有参数设置和威胁模型中都是一致的。最后，我们研究了洗牌过程的两种常见变体，这些变体会导致进一步的隐私泄漏（高达10倍）。总体而言，我们的工作强调了在没有严格分析方法的情况下使用洗牌而非泊松子采样的风险。</span></span></p><p cid="n973" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f597-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f597-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n975" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">236、Token Time Bomb: Evaluating JWT Implementations for Vulnerability Discovery</span></span></p><p cid="n976" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">JSON Web令牌（JWT）已成为现代分布式Web应用中安全信息交换的广泛采用标准，特别是在身份验证和授权场景中。然而，JWT的实现引入了各种漏洞，例如签名验证绕过、令牌欺骗和拒绝服务攻击。尽管先前研究已报告了此类个别漏洞，但缺乏对JWT实现的系统性研究。在本文中，我们提出了JWTFuzz，一种新颖的测试方法，用于有效发现JWT实现中的漏洞。我们对10种流行编程语言中的43个JWT实现进行了JWTFuzz评估，发现了31个先前未知的安全漏洞，其中20个已被分配CVE编号。我们展示了这些漏洞的安全影响，例如在Kubernetes中实现身份验证绕过和对Apache James的拒绝服务攻击。我们进一步将这些漏洞分为五类，并提出了几种缓解策略。我们与国际互联网工程任务组（IETF）讨论了我们的缓解策略，他们已认可我们的发现，并建议在新RFC文档中采用我们的缓解措施。我们还向相关提供商报告了已识别的漏洞，并收到了Apache、Connect2id、Kubernetes、Let&#39;s Encrypt和RedHat的确认和漏洞赏金奖励。</span></span></p><p cid="n977" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f697-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f697-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n979" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">237、Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models</span></span></p><p cid="n981" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">以DALL·E和Midjourney为代表的文本到图像（T2I）模型因能创建逼真的图像而广受欢迎。这些图像的质量依赖于精心设计的提示词，这些提示词已成为宝贵的知识产权。虽然熟练的提示词创作者在市场上展示其AI生成的艺术作品以吸引买家，但这种业务无意使他们面临&#34;提示词窃取攻击&#34;。现有的最先进攻击技术通过针对特定模型的训练，从固定的修饰符集合（即风格描述）中重建提示词，这些技术在适应不同展示作品（即目标图像）和扩散模型方面表现出有限的适应性和有效性。为缓解这些限制，我们提出了Prometheus，一种无需训练、包含中间代理、基于搜索的提示词窃取攻击方法，通过与本地代理模型交互来逆向工程展示作品中的宝贵提示词。该方法包含三项创新设计。首先，我们引入了动态修饰符，作为先前工作中使用的静态修饰符的补充。这些动态修饰符提供了更多与展示作品相关的具体细节，我们利用自然语言处理分析即时生成它们。其次，我们设计了一种上下文匹配算法，用于对动态和静态修饰符进行排序。这一离线过程有助于减少后续步骤的搜索空间。第三，我们与本地代理模型交互，使用贪心搜索算法逆向提示词。基于反馈指导，我们优化提示词以实现更高的保真度。评估结果显示，Prometheus成功地从PromptBase和AIFrog等流行平台提取提示词，针对Midjourney、Leonardo.ai和DALL·E等多样化的受害者模型，实现了25.0%的攻击成功率提升。我们还验证了Prometheus能够抵抗广泛的潜在防御措施，进一步突显了其在实践中的严重性。</span></span></p><p cid="n982" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1899-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1899-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n984" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">238、TranSPArent: Taint-style Vulnerability Detection in Generic Single Page Applications through Automated Framework Abstraction</span></span></p><p cid="n985" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">单页应用（SPA）框架允许开发者使用高级组件（如搜索框）在单个HTML页面中构建复杂的Web应用程序。SPAs面临的一个研究问题是如何检测污点式漏洞，因为SPA框架以新形式重新引入了不安全的DOM API，例如将SPA组件参数作为污点汇点。尽管先前的研究已致力于改进SPAs中的漏洞检测，但据我们所知，这些方法严重依赖硬编码的污点汇点，这不仅需要针对不同的SPA框架进行手动维护，还可能遗漏某些不安全的SPA API，从而导致检测到的漏洞出现漏报。在本文中，我们提出了TranSPArent，一个SPA漏洞检测工具，它通过结合静态分析和动态分析自动抽象SPA框架，以揭示框架特定的汇点，从而促进端到端的静态漏洞检测。TranSPArent首先从不安全的DOM API列表执行反向污点分析，直到框架接口，以揭示接口的哪些部分可能污染DOM API。这种自动框架抽象每个SPA框架只需执行一次。然后，TranSPArent检测发现的SPA汇点与攻击者可控源之间的数据流路径，以检测每个应用程序中的污点式漏洞。我们在GitHub仓库数据库上评估了TranSPArent，发现了11个零日漏洞，包括一个拥有24k+ GitHub星标和每月3000万请求的仓库。迄今为止，其中四个零日漏洞已被开发者修复和/或确认。在我们的评估过程中，TranSPArent从三种最广泛使用的SPA框架（Vue、React和Angular）中总共发现了19个中间SPA汇点。其中14个新发现的汇点未在CodeQL标准库（最先进的静态分析工具）中列出。</span></span></p><p cid="n986" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1721-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1721-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n988" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">239、Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias</span></span></p><p cid="n990" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型(LLMs)越来越多地被用于执行大规模的自动化代码审查和静态分析，支持漏洞检测、代码摘要和重构等任务。在本文中，我们识别并利用了基于LLM的代码分析中的一个关键漏洞：一种抽象偏差，导致模型过度泛化熟悉的编程模式而忽略微小但有意义的错误。攻击者可以利用这个盲点，通过最小程度的修改来劫持LLM的解释控制流，同时不影响实际的运行时行为。我们将这种攻击称为熟悉模式攻击(FPA)。我们开发了一个全自动的黑盒算法，用于发现并向目标代码中注入FPA。我们的评估表明，FPA不仅对基础模型和推理模型有效，而且可以在不同模型家族(OpenAI、Anthropic、Google)之间迁移，并且在多种编程语言(Python、C、Rust、Go)中具有通用性。此外，即使模型通过强大的系统提示被明确警告了这种攻击，FPA仍然有效。最后，我们探讨了FPA的积极防御用途，并讨论了它们对面向代码的LLM可靠性和安全的更广泛影响。</span></span></p><p cid="n991" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2066-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2066-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n993" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">240、UIEE: Secure and Efficient User-space Isolated Execution Environment for Embedded TEE Systems</span></span></p><p cid="n994" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可信执行环境（TEE）已被广泛探索用于增强嵌入式系统的安全性。现有的嵌入式TEE系统运行时占用较小的内存空间，仅提供安全关键功能，以保持最小的可信计算基（TCB）。不幸的是，这种设计选择导致这些TEE系统软件资源不足，难以在嵌入式TEE内执行具有大型代码库的复杂应用程序。在本文中，我们提出了一种用户空间隔离执行环境（UIEE），通过在TEE内直接运行未经修改的数据处理应用程序来增强TEE功能，同时不增加TCB大小。UIEE通过为应用程序动态分配足够的内存区域来构建沙箱环境，并将其与丰富执行环境（REE）和TEE隔离，从而保护UIEE免受REE攻击，同时保护TEE免受潜在受损的UIEE应用程序的侵害。此外，我们提出了一种基于库操作系统（即Linux内核库，LKL）的UIEE运行时环境，可为UIEE应用程序提供标准C运行时API。为了解决LKL的并发问题，我们提出了一种LKL线程同步机制，在具有单线程执行模型的UIEE内运行多线程LKL。此外，我们还设计了一种新颖的按需线程迁移机制，以实现在UIEE内的LKL上下文切换。我们在NXP IMX6Q SABRE-SD评估板上实现并部署了一个UIEE原型，成功在UIEE内运行了8个未经修改的真实世界基于libc的应用程序。实验结果表明，UIEE带来的性能开销可以忽略不计。我们是第一个提出面向TrustZone的LibOS并评估其可行性和安全特性的研究。</span></span></p><p cid="n995" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s208-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s208-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n997" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">241、Understanding the Status and Strategies of the Code Signing Abuse Ecosystem</span></span></p><p cid="n999" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">使用数字证书对软件签名是其可信性和完整性的重要保障。然而，攻击者可以滥用这一机制为恶意样本获取签名，从而促进恶意软件的传播。尽管已有工作揭示了代码签名滥用的实例，但这一问题仍然存在且不断升级。理解生态系统的演变和滥用者的策略对于改进防御机制至关重要。在本工作中，我们对代码签名滥用进行了大规模测量，使用了从野外收集的3,216,113个已签名的恶意PE文件。通过细粒度分类，我们识别出43,286个被滥用的证书，并将其分为五种滥用类型，创建了迄今为止最大的标记数据集。我们的分析表明，滥用仍然普遍存在，涉及来自114个国家、由46个证书颁发机构（CA）发行的证书。我们还观察到了滥用者技术的演变，并识别了证书撤销方面的当前局限性。此外，我们表征了滥用者的行为和策略，揭示了五种规避检测、降低成本和增强滥用影响的策略。值得注意的是，我们发现了3,484个多态证书集群，并首次记录了恶意软件利用多态技术规避撤销检查的实际案例。我们的研究结果揭示了当前代码签名实践中的关键缺陷，预计将提高社区对滥用威胁的认识。</span></span></p><p cid="n1000" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2857-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2857-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1002" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">242、Understanding the Stealthy BGP Hijacking Risk in the ROV Era</span></span></p><p cid="n1003" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">路由源验证（ROV）的部分部署带来了一种意外的安全威胁，称为隐蔽的BGP劫持，即一种特别难以察觉的BGP劫持形式，其中恶意路由可以在不到达（从而提醒）受害者的情况下转移流量。这一风险在很大程度上仍未被探索，既没有记录在案的现实世界事件，也没有系统性的特征描述。为了填补这一空白，我们形式化了隐蔽的BGP劫持，并提出了启发式方法，通过路由表差异来发现潜在实例。我们进行了首次实证研究，以跟踪和描述现实世界中的隐蔽BGP劫持，贡献了一个精选的现实世界事件数据集和一个长期监控服务。受实证见解的启发，我们进一步进行了分析研究，以全面评估风险。这需要准确的ROV部署数据、完整的全球互联网路由以及定制的分析模型。为了应对这些挑战，我们开发了SHAMAN，一个专门用于评估隐蔽BGP劫持风险的BGP路由推断框架。SHAMAN整合多种来源构建准确的ROV部署视图，通过高效的基于矩阵的方法推断完整的全球互联网路由，并通过&#34;受害者-目标-劫持者&#34;三元组模型促进统计风险分析。SHAMAN将生成互联网规模路由的时间从三个月以上缩短到仅5.22小时，使得在现实ROV部署下能够对83亿条生成路由进行系统性风险评估。我们的研究结果显示隐蔽BGP劫持的总体成功概率为14.1%，而在特定情况下，有针对性的攻击成功率高达99.5%。与我们现实世界数据集的验证显示，事件级别的准确度高达95.9%，证明了我们分析结果的真实性。</span></span></p><p cid="n1004" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s97-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s97-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1006" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">243、Unknown Target: Uncovering and Detecting Novel In-Flight Attacks to Collision Avoidance (TCAS)</span></span></p><p cid="n1007" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">是防止空中碰撞的强制性最后安全保障。尽管该系统具有关键的安全作用，但其未经认证和加密的通信协议长期以来一直被确认为安全风险。尽管研究人员先前已经展示了实际的注入攻击，但官方评估认为这些漏洞仅限于实验室环境，并指出目前尚无缓解措施。在本文中，我们对这两种说法提出质疑。我们提供了有力证据表明，针对TCAS的飞行中网络攻击已经发生。通过对一系列涉及多架飞机的异常事件的公开飞行和通信数据进行详细分析，我们确定了一种与幽灵飞机注入攻击一致的独特特征。我们详细说明了这种新型攻击如何利用传统协议特性，并描述了三种复杂度递增的攻击策略；其中最具攻击性的策略可以将目标的感知距离减少3.5公里以上，足以从远距离触发受害飞机的防撞警报。我们实现了与观察到的事件最一致的攻击策略，并进行了实验评估，实现了1.9公里的欺骗性距离减少，证实了其可行性。此外，为应对此类威胁提供基础，我们提出了一种新颖的、向后兼容的方法，通过重新利用受害者广播的TCAS警报数据来地理定位此类攻击的来源。在最可能的攻击变体模拟场景中，我们的方法实现了855米的中位数定位精度。将此技术应用于真实事件数据，我们能够识别出异常现象以及观察到的幽灵飞机注入攻击的可能来源。</span></span></p><p cid="n1008" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1806-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1806-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1010" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">244、Unshaken by Weak Embedding: Robust Probabilistic Watermarking for Dataset Copyright Protection</span></span></p><p cid="n1012" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在现代数据即服务（DaaS）生态系统中，数据策展商（如数据经纪公司）从众多贡献者处聚合高质量数据，并为深度学习模型提供商将其变现。然而，恶意策展商可能出售有价值的数据却不告知其原始贡献者，这违反了个人利益和法律。侵入式水印是保护数据版权的最先进技术之一，它能检测可疑模型是否携带预定义模式。然而，这些方法面临诸多限制：在低水印注入率（≤1.0%）下难以工作；性能下降；误报；对水印清洗缺乏鲁棒性。</span></span></p><p cid="n1013" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">本文提出了一种创新的侵入式水印方法，称为</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DIP</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">（数据智能概率水印），以支持数据集所有权验证，同时解决上述局限性。它应用了感知分布的样本选择算法，嵌入带水印样本与多个输出之间的概率关联，并采用双重验证框架，同时利用推理结果及其分布作为水印信号。在4个图像和5个文本数据集上的广泛实验表明，</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DIP</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">保持了模型性能，并在1%的注入预算下实现了89.4%的平均水印成功率。我们进一步验证了</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DIP</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">与各种水印数据设计正交，并能无缝整合其优势。此外，</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DIP</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在多种模态（图像和文本）和任务（回归）上证明有效，在大语言模型的生成任务上也表现出色。</span></span><span md-inline="em" style="box-sizing: border-box;"><em style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">DIP</span></span></em></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">对各种对抗环境具有鲁棒性，包括3种基于数据增强、3种基于数据清洗、4种基于鲁棒训练和3种基于合谋的水印移除方法，而现有的最先进方法则无法应对。源代码已发布于</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://github.com/SixLab6/DIP" target="_blank">https://github.com/SixLab6/DIP</a></span></span><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">。</span></span></p><p cid="n1014" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1356-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1356-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1016" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">245、Unveiling BYOVD Threats: Malware’s Use and Abuse of Kernel Drivers</span></span></p><p cid="n1017" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">&#34;自带漏洞驱动程序&#34;（BYOVD）攻击利用合法的、经过数字签名的Windows驱动程序中隐藏的缺陷，使攻击者能够进入内核空间，禁用安全控制，并执行从勒索软件到国家支持的网络间谍活动的隐蔽行动。由于大多数公共沙箱仅检查用户模式活动，这种内核级别的滥用通常难以被发现。在这项工作中，我们首先介绍了首个BYOVD行为的动态分类法。该分类法基于对实际事件的手动调查和细粒度内核跟踪分析综合而成，将每次攻击映射到连续的阶段，并列举了每个步骤中被滥用的关键API。然后，我们提出了一种基于虚拟化的沙箱，它跟踪驱动程序执行路径的每一步，从最初的用户模式请求到最低级别的内核指令，而无需重新签名驱动程序或修改主机。最后，沙箱自动为每个观察到的动作添加相应的分类注释，生成一份分阶段报告，突出显示样本表现出可疑行为的位置和方式。针对当前的BYOVD技术环境进行测试，我们分析了8,779个加载了773个不同签名驱动程序的恶意软件样本。该沙箱标记了48个驱动程序的可疑行为，随后的手动验证导致向微软、其供应商和公共威胁情报平台披露了七个先前未知的漏洞驱动程序。我们的结果表明，对内核控制流的深入、透明跟踪可以揭示传统分析流程无法发现的BYOVD滥用行为，丰富了社区对驱动程序利用的知识，并促进了Windows防御的主动加固。</span></span></p><p cid="n1018" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1491-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1491-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1020" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">246、User-Space Dependency-Aware Rehosting for Linux-Based Firmware Binaries</span></span></p><p cid="n1021" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">固件重托管是一种基础仿真技术，能够大规模地对固件二进制文件进行动态分析。成功重托管基于Linux的固件服务需要正确模拟系统级功能（如设备接口）和用户空间依赖项（如配置文件、进程间通信）。然而，现有解决方案未能充分利用用户空间知识。作为第一个用户空间进程的初始化例程负责设置操作环境，但往往执行不完整，导致初始化不完整。此外，所有仿真故障都被统一处理，无法区分系统级仿真问题及其对用户空间依赖项的间接影响。为填补这一空白，我们开发了FIRMWELL框架，该框架首先将固件重托管建模为目标二进制文件及其用户空间依赖项的协同仿真。它首先重托管初始化例程以构建环境，然后启动目标服务，这一过程通常涉及一百多个进程。当仿真故障发生时，FIRMWELL会识别阻塞进程，分析错误仿真的资源，并应用有针对性的修复。关键策略是通过纠正底层系统级仿真错误来解决用户空间依赖项故障，同时利用程序分析进行精确的资源值推断。在对14,049个固件镜像的评估中，FIRMWELL成功重托管了6,490个服务，比现有最佳方法高出1.6-8倍（FirmAE为3,581个，Greenhouse为3,962个，Pandawan为810个），同时将平均重托管时间减少了1.8-8.4倍（分别为12分钟、22分钟、74分钟和101分钟）。FIRMWELL被应用于对1,043个固件镜像进行模糊测试，发现了67个零日漏洞，其中10个已被分配CVE标识符。</span></span></p><p cid="n1022" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s249-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s249-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1024" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">247、Validity Is Not Enough: Uncovering the Security Pitfall in Chainlink’s Off-Chain Reporting Protocol</span></span></p><p cid="n1025" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">区块链预言机在将链外交易所的价格数据传递给智能合约方面发挥着关键作用，从而实现自动化金融服务。作为主导的预言机服务提供商，Chainlink采用去中心化预言机网络(DON)来提供价格数据。在Chainlink的DON中，多个预言机节点独立观察加密货币的价格并运行链下报告(OCR)协议，从它们的观测值中确定一个唯一价格。源自OCR协议的价格偏差将带来安全风险。为防止拜占庭预言机节点引发任意价格偏差，OCR的有效性属性确保确定的价格被诚实观测值所限制。然而，这一界限在实际环境中仍不明确，且拜占庭行为仍能引发多大程度的价格偏差尚不清楚。本文通过实证和理论分析，深入研究了拜占庭行为对OCR协议中确定价格的潜在影响。首先，我们的实证分析显示，在实际环境中，拜占庭行为在OCR协议中仍有足够空间影响确定的价格。随后，我们详细阐述了战略性地影响确定价格的拜占庭行为，并对其影响进行了形式化建模。此外，我们使用Chainlink的真实世界价格数据评估了这些拜占庭行为的影响。实验结果表明，拜占庭行为引发的价格偏差可达ETH价格的8.47%。我们的案例研究进一步表明，被拜占庭行为影响的价格值可能带来下游金融影响，规模可达10^5美元，而此类价格值的累积影响可能达到数百万美元。总之，这项工作揭示出，即使在有效性保证下，拜占庭行为仍可能对OCR协议中的确定价格产生不可忽视的影响。我们已将发现结果向Chainlink进行了道德报告，旨在支持OCR协议的安全性。</span></span></p><p cid="n1026" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f458-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f458-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1028" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">248、Vault Raider: Stealthy UI-based Attacks Against Password Managers in Desktop Environments</span></span></p><p cid="n1029" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">密码管理器通过生成强大且唯一的密码显著改进了基于密码的身份验证，同时通过自动填充功能简化了实际的身份验证过程。关键的是，在传统浏览环境中使用时，自动填充提供了额外的安全保护，因为它可以轻易地挫败网络钓鱼攻击，因为网站域名信息唾手可得。随着主要网络服务越来越多地部署独立的原生应用程序的趋势，密码管理器也开始为桌面环境提供通用自动填充和其他用户友好的功能。然而，目前尚不清楚密码管理器的安全保护在这些环境中如何运作。在本文中，我们通过首次对流行密码管理器（包括1Password和LastPass）在主要桌面环境（macOS、Windows、Linux）中提供的自动填充相关功能进行系统性实证分析，填补了这一空白。我们通过实验发现，密码管理器采用不同的策略与桌面应用程序交互，并采用不同级别的针对基于用户界面攻击的保护措施。例如，在macOS上，我们发现可以利用操作系统提供的API和检查实现高级别的安全性，而在Windows上，我们识别出缺乏适当的安全检查，这主要是由于操作系统的限制。在每种情况下，我们都展示了概念验证攻击，这些攻击允许其他应用程序绕过现有的安全检查，并通过不可见的模拟按键 stealthily 窃取用户的凭据、一次性密码和保险库密钥。因此，我们提出了一系列可以缓解我们攻击的对策。由于我们攻击的严重性，我们向被分析的密码管理器供应商披露了我们的发现和建议的对策，这已经促使某些供应商开始修复过程，并获得了错误赏金。最后，我们将分享我们的代码，以促进加强密码管理器的额外研究。</span></span></p><p cid="n1030" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1067-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1067-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1032" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">249、VDORAM: Towards a Random Access Machine with Both Public Verifiability and Distributed Obliviousness</span></span></p><p cid="n1033" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">可验证随机访问机（vRAM）作为一种基础模型，能够表达具有可证明安全保证的复杂计算，应用于安全电子投票、金融审计和隐私保护智能合约等领域。然而，现有的vRAM均未提供分布式 obliviousness，这在多个证明者希望防止彼此之间以及与验证者之间信息泄露的场景中是一个关键需求，因为现有解决方案难以解决MPC与ZKP之间的范式不匹配问题，这限制了实际多证明者ZKP前端的发展。这一差距的出现是因为MPC协议针对最小计算进行了优化，而ZKPs需要完整的计算轨迹用于证明。此外，调整RAM设计也面临挑战，因为vRAM并非为盲目执行的高成本而设计，且现有的DORAM缺乏公开可验证性。为应对这些挑战，我们引入了CompatCircuit，据我们所知，这是首个多证明者ZKP前端实现，旨在弥合这一差距。CompatCircuit将协作zkSNARKs与新型MPC协议相结合，将计算和验证统一为单个兼容的电路范式。基于CompatCircuit，我们提出了VDORAM，这是首个公开可验证的分布式 oblivious RAM。VDORAM平衡了在线MPC的高通信延迟与离线证明生成的复杂性，形成了一种能够兼顾这些竞争性需求的RAM设计。我们用约15,000行代码实现了CompatCircuit和VDORAM，通过微基准测试、比较分析和程序示例等大量实验证明了它们的实际可行性。</span></span></p><p cid="n1034" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s16-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s16-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1036" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">250、VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs</span></span></p><p cid="n1037" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">微调大型语言模型（LLMs）对于使其适应特定任务至关重要，但这仍然计算密集，并且在不可信环境中引发了正确性和隐私方面的担忧。尽管像低秩适应（LoRA）这样的参数高效方法显著降低了资源需求，但在零知识约束下确保微调的安全性和可验证性仍然是一个未解决的挑战。为此，我们提出了VeriLoRA，这是第一个将LoRA微调与零知识证明（ZKPs）相结合的框架，实现了可证明的安全性和正确性。VeriLoRA采用先进的密码学技术——如查找参数、求和协议和多项式承诺——来验证基于Transformer架构中的算术和非算术操作。该框架为LoRA微调过程中的前向传播、反向传播和参数更新提供端到端的可验证性，同时保护模型参数和训练数据的隐私。基于GPU的实现，VeriLoRA在开源LLMs（如LLaMA）上的实验验证中展示了其实用性和效率，可扩展至130亿参数。通过将参数高效微调与ZKPs相结合，VeriLo弥合了一个关键差距，使LLMs能够在敏感或不可信环境中安全可信地部署。</span></span></p><p cid="n1038" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2361-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2361-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1040" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">251、VICTOR: Dataset Copyright Auditing in Video Recognition Systems</span></span></p><p cid="n1041" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">视频识别系统正日益应用于日常生活，如内容推荐和安全监控。为促进视频识别的发展，许多机构已发布了高质量的开源许可公共数据集，用于训练先进模型。同时，这些数据集也容易被滥用和侵权。数据集版权审核是识别此类未经授权使用的有效解决方案。然而，现有的数据集版权解决方案主要关注图像领域；视频数据的复杂性使得视频领域的数据集版权审核尚未得到探索。具体而言，视频数据引入了额外的时间维度，这对现有方法的有效性和隐蔽性构成了重大挑战。在本文中，我们提出了VICTOR，这是首个面向视频识别系统的数据集版权审核方法。我们开发了一种通用且隐蔽的样本修改策略，能够增强目标模型的输出差异。通过仅修改少量样本（例如1%），VICTOR放大了已发布修改样本对目标模型预测行为的影响。然后，模型对已发布修改样本和未发布原始样本的行为差异可作为数据集审核的关键依据。在多个模型和数据集上的广泛实验凸显了VICTOR的优越性。最后，我们证明了在面对针对训练视频或目标模型的多种干扰机制时，VICTOR具有鲁棒性。</span></span></p><p cid="n1042" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f746-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f746-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1044" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">252、ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks</span></span></p><p cid="n1045" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">深度伪造技术的迅速崛起产生了逼真但虚假的数字内容，威胁着媒体的真实性。深度伪造技术操纵视频、图像和音频，传播错误信息，模糊真实与虚假的界限，凸显了对有效检测方法的需求。传统的深度伪造检测方法往往难以应对复杂、定制的深度伪造内容，特别是在泛化能力和对抗恶意攻击的鲁棒性方面。本文介绍了ViGText，一种创新方法，它将图像与基于视觉的大语言模型(VLLM)文本解释在基于图的框架中集成，以改进深度伪造检测。ViGText的创新之处在于它将详细解释与视觉数据相结合，提供了比通常缺乏特异性且无法揭示细微不一致性的字幕更具上下文感知能力的分析。ViGText系统地将图像分割为块，构建图像和文本图，并利用图神经网络(GNN)进行集成分析以识别深度伪造。通过在空间和频域进行多级特征提取，ViGText捕捉了增强其鲁棒性和准确性的细节，能够检测复杂的深度伪造内容。大量实验表明，ViGText显著提高了泛化能力，并在检测用户定制的深度伪造内容时取得了显著的性能提升。具体而言，在泛化评估中，平均F1分数从72.45%上升到98.32%，反映了该模型对未见过的、经过微调的稳定扩散模型变体的优越泛化能力。在鲁棒性方面，ViGText与其他深度伪造检测方法相比，在面对最先进的基础模型对抗攻击时，召回率提高了11.1%。ViGText在面对利用其图架构的针对性攻击时，将分类性能的降低限制在4%以内，同时略微增加了执行成本。ViGText将细粒度的视觉分析与文本解释相结合，为深度伪造检测建立了新的基准，并为保护媒体真实性和信息完整性提供了更可靠的框架。</span></span></p><p cid="n1046" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s303-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s303-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1048" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">253、vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis</span></span></p><p cid="n1049" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">二进制代码相似性分析（BCSA）在许多安全任务中发挥着至关重要的作用，包括恶意软件分析、漏洞检测和软件供应链安全。尽管过去十年提出了许多BCSA技术，但很少有利用寄存器和内存值的语义进行比较的，尽管初步结果很有前景。现有的基于值的方法通常仅关注在编译设置中保持不变的值，从而忽略了更广泛的语义丰富信息。在本文中，我们确定了限制基于值的BCSA有效性的三个核心挑战：值提取的可扩展性不足、缺乏噪声过滤以及值比较效率低下。这些缺点既限制了语义覆盖范围，也影响了可扩展性。为了充分释放基于值的BCSA的潜力，我们提出了vSim，这是一个新颖的框架，能够系统性地捕获所有寄存器和内存操作的值，过滤掉语义无关的值（例如全局地址），并对剩余值进行归一化和传播，以实现健壮且可扩展的相似性分析。广泛的评估表明，vSim在准确性、鲁棒性和可扩展性方面 consistently优于最先进的BCSA系统。它在不同架构和工具链上具有良好的泛化能力，能够在多样化的数据集上产生可靠的结果。</span></span></p><p cid="n1050" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f213-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f213-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1052" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">254、VulSCA: A Community-Level SCA Approach for Accurate C/C++ Supply Chain Vulnerability Analysis</span></span></p><p cid="n1053" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着第三方库（TPLs）在C/C++开发中的广泛采用，软件供应链安全变得至关重要。现有的C/C++供应链漏洞分析方法存在显著局限性。一些方法仅专注于依赖识别，导致误报（FPs），而另一些方法强调漏洞检测却忽略依赖关系，需要耗时的完整仓库扫描，从而阻碍了对供应链漏洞的快速响应。为此，我们探讨了准确依赖构建和漏洞检测的适当粒度。我们提出了一种社区级的软件成分分析（SCA）方法，将项目的调用图建模为社会网络并应用社区检测。然后通过社区相似性建立项目与TPLs之间的依赖关系。对于漏洞检测，我们在依赖社区内执行基于克隆的检测以验证漏洞的存在，并引入两阶段可达性分析以确定这些漏洞是否可以传播到目标项目。我们实现了VulSCA，这是首个集成漏洞检测和可达性分析的C/C++ SCA框架。实验结果表明，在SCA方面，VulSCA的性能优于CENTRIS和OSSFP，F1-score提高了4-12%。在供应链漏洞检测方面，其F1-score比基于版本的方法高44-48%，比基于代码的方法高17-23%。在效率方面，VulSCA的整体开销低于所有基于代码的方法。此外，VulSCA在广泛使用的开源项目中识别出32个先前未被修复的供应链漏洞，这些漏洞已报告给相应的供应商。</span></span></p><p cid="n1054" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s613-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s613-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1056" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">255、Was My Data Used for Training? Membership Inference in Open-Source LLMs via Neural Activations</span></span></p><p cid="n1057" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着大型语言模型（LLMs）的快速发展，其应用已扩展到日常生活的各个方面。开源LLMs因其可访问性而广受欢迎，导致广泛下载和再分发。LLMs的强大能力源于对大规模且通常未公开的数据集的训练，这引发了关于是否包含版权或个人数据等敏感内容的问题，这被称为成员推断问题。现有方法主要依赖模型输出，而忽略了丰富的内部表示。内部数据的有限访问导致次优结果，揭示了开源白盒LLMs中成员推断的研究空白。在本文中，我们解决了检测开源LLMs训练数据的挑战。为支持这项研究，我们引入了三个动态基准：WikiTection、NewsTection和ArXivTection。随后，我们提出了一种通过分析LLMs的神经激活来进行训练数据检测的白盒方法。我们的关键见解是，LLMs所有层的神经元激活反映了输入数据在LLM内部的知识表示，能够有效区分LLM的训练数据和非训练数据。在这些基准上的广泛实验证明了我们方法的有效性。例如，在WikiTection基准上，我们的方法在五个LLMs（GPT2-xl、LLaMA2-7B、LLaMA3-8B、Mistral-7B和LLaMA2-13B）上均实现了约0.98的AUC。此外，我们对模型大小、输入长度和文本改写等因素进行了深入分析，进一步验证了我们方法的鲁棒性和适应性。</span></span></p><p cid="n1058" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f474-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f474-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1060" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">256、WBSLT: A Framework for White-Box Encryption Based on Substitution-Linear Transformation Ciphers</span></span></p><p cid="n1061" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">加密算法面临各种密钥提取攻击，促使在不同威胁模型下产生多种防御工作。其中，白盒威胁模型具有最强的对抗场景，攻击者可以完全访问和控制密码学实现及其执行环境。然而，先前的白盒加密设计主要保护单个密钥相关表，使得白盒和侧信道攻击能够恢复密钥。基于我们的观察，对这些表的边界进行模糊化可以使攻击无效。因此，我们提出了WBSLT，一种用于替换-线性变换（SLT）密码表格式白盒实现的新型设计框架。WBSLT通过线性和非线性变换保护嵌入密钥的表，并部分将每个组件的计算留给下一个组件，以减轻单个密钥相关表泄露。为进一步防御差分计算分析和差分故障分析，该框架集成了掩码、随机化和外部编码。理论分析表明其对各种攻击具有免疫性。实验结果验证了WBSLT在多个计算平台上的实用性，显示出高效的加密性能和合理的内存消耗。</span></span></p><p cid="n1062" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2492-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2492-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1064" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">257、WCDCAnalyzer: Scalable Security Analysis of Wi-Fi Certified Device Connectivity Protocols</span></span></p><p cid="n1065" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">Wi-Fi联盟已开发了多种设备连接协议，如Wi-Fi Direct、Wi-Fi EasyConnect和Wi-Fi EasyMesh，这些协议在全球数十亿设备中发挥着关键作用。鉴于其广泛采用，确保这些协议的安全性和隐私性至关重要。然而，现有研究尚未全面审视这些协议设计的安全性和隐私性方面。为填补这一空白，我们推出了WCDCAnalyzer（Wi-Fi认证设备连接分析器），这是一个正式分析框架，旨在评估这些广泛使用的Wi-Fi认证设备连接协议的安全性和隐私性。在形式化验证Wi-Fi Direct协议时，一个重大挑战是由协议规模大、复杂性高导致的状态爆炸问题所引起的可扩展性问题，这会导致内存使用呈指数级增长。为应对这一挑战，我们开发了一种遵循组合推理范式的系统分解方法，并将其整合到WCDCAnalyzer中。这使得WCDCAnalyzer能够自动将给定协议分解为多个子协议，分别验证每个子协议，然后合并结果。我们的设计是基于严格基础的组合推理的实际应用，我们提供了详细算法，展示了如何将这种推理方法应用于密码协议验证。使用WCDCAnalyzer，我们分析了这些协议并发现了10个漏洞，包括身份验证绕过、隐私泄露和拒绝服务攻击。这些漏洞及相关实际攻击已在商业设备上得到验证，并获得了Wi-Fi联盟的认可。</span></span></p><p cid="n1066" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1049-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1049-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1068" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">258、What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs</span></span></p><p cid="n1070" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">开源软件项目是现代软件生态系统的基础，Linux内核因其普遍性和复杂性而成为关键典范。尽管安全补丁持续集成到Linux主线内核中，但下游维护者常常延迟采用这些补丁，从而造成漏洞窗口。这种滞后的一个关键原因是难以识别安全关键补丁，特别是那些处理可利用漏洞的补丁，如越界（OOB）访问和使用后释放（UAF）错误。由于故意静默的bug修复、不完整或缺失的CVE分配、CVE发布延迟以及最近Linux内核CVE分配标准的变更，这一挑战进一步加剧。以往的工作如GraphSPD提出了二元分类器来区分安全补丁与非安全补丁。然而，这些方法不能提供漏洞类型的细粒度分类，这对于优先修复OOB和UAF等高影响错误至关重要。我们的工作旨在将这些粗略标记的安全补丁分类为细粒度类别，即OOB、UAF或非OOB-UAF类型。尽管细粒度补丁分类方法已经存在，但它们在覆盖范围和准确性方面都存在局限性。在这项工作中，我们确定了以前未被探索的机会，可以显著改进细粒度补丁分类。具体而言，通过利用提交标题/消息和差异的线索以及适当的代码上下文，我们开发了DUALLM，这是一个双方法流水线，集成了基于大型语言模型（LLM）和微调小型语言模型的两种方法。DUALLM实现了87.4%的准确率和0.875的F1分数，显著优于先前解决方案。值得注意的是，DUALLM成功地将5,140个最近的Linux内核补丁中的111个识别为处理OOB或UAF漏洞，其中90个真阳性通过手动验证确认（许多在补丁描述中没有明确指示）。此外，我们为两个已识别的错误（一个UAF和一个OOB）构建了概念验证，包括一个用于执行先前未知的控制流劫持的漏洞，进一步证明了分类的正确性。</span></span></p><p cid="n1071" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s328-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s328-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1073" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">259、When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its Countermeasures</span></span></p><p cid="n1074" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">大型语言模型（LLM）的出现催生了广泛应用，包括代码生成、聊天机器人和AI智能体。然而，部署这些应用在成本和效率方面面临重大挑战。应对这些挑战的一种重要优化是语义缓存，它基于语义相似性跨用户重用查询-响应对。这种机制在学术界和工业界都获得了广泛关注，并已被集成到Azure、AWS和阿里巴巴等云服务提供商的LLM服务基础设施中。本文首次证明语义缓存容易受到缓存投毒攻击，即攻击者注入精心设计的缓存条目，导致其他用户接收到攻击者定义的响应。我们在多种场景下演示了语义缓存投毒攻击，并确认其在三大主要公有云中的实用性。基于这些攻击，我们评估了现有的对抗性提示防御方法，发现它们对语义缓存投毒无效，促使我们提出了一种新的防御机制，相比现有方法显示出更好的保护效果，尽管完全缓解仍然具有挑战性。我们的研究表明，缓存投毒这一长期存在的安全问题在LLM系统中重新出现。虽然我们的分析聚焦于语义缓存，但潜在风险可能延伸至LLM系统中使用的其他类型的缓存机制。</span></span></p><p cid="n1075" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f200-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f200-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1077" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">260、When Focus Enhances Utility: Target Range LDP Frequency Estimation and Unknown Item Discovery</span></span></p><p cid="n1078" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">局部差分隐私（LDP）协议能够收集随机化的客户端消息用于数据分析，而无需可信的数据管理员。此类协议已被谷歌、苹果和微软等大型科技公司成功应用于实际场景。在本文中，我们提出了一种广义计数均值草图（GCMS）协议，该协议涵盖了多种现有的频率估计协议。我们的方法显著改善了通信、隐私和准确性之间的三向权衡。我们还引入了一种通用效用分析框架，能够优化参数设计。基于此，我们提出了一种最优计数均值草图（OCMS）框架，用于最小化收集具有目标频率项目的方差。此外，我们提出了一种用于收集未知领域数据的新协议，因为我们的频率估计协议仅对已知数据领域有效。利用基于稳定性的直方图技术与加密-打乱-分析（ESA）框架相结合，我们的方法采用辅助服务器构建直方图，而无需访问原始数据消息。该协议实现了与中心DP模型相当的准确性，同时提供了类本地隐私保证，并显著降低了计算成本。</span></span></p><p cid="n1079" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s1397-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s1397-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1081" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">261、When Mixnets Fail: Evaluating, Quantifying, and Mitigating the Impact of Adversarial Nodes in Mix Networks</span></span></p><p cid="n1082" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">混合网络（mixnets）通过将数据包独立地随机经过选定的跳节点（mixnodes）传输，为客户端提供针对强大网络对手的通信匿名性，从而破坏数据包的可链接性。尽管这种方法在Nym系统中实现，能够最大化对网络对手的混淆效果，但它允许攻击者通过控制部分mixnodes（节点总数的10%/5%）来完全消除通信量与其目的地超过特定阈值（4MB/30MB）的所有客户端的匿名性。为缓解此类漏洞，本研究开发了一系列新颖的路径选择技术，实现了对网络对手的抵抗力与对受损mixnodes的弹性之间的平衡。鉴于现有的匿名性度量不足以量化混合网络中的 adversarial 风险，我们额外引入了有效的基于实证和模拟的度量指标。通过理论、实证和基于模拟的评估，我们全面评估了所提出的方案，证明这些方法可将对受损节点的脆弱性降低高达80%，同时为网络对手带来的有限优势。我们的分析进一步揭示，最先进的匿名性度量指标与我们所提出的度量指标相比，会产生误导性结果，这些结果影响了Nym系统中的某些设计决策。</span></span></p><p cid="n1083" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2384-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2384-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1085" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">262、WhiteCloak: How to Hold Anonymous Malicious Clients Accountable in Secure Aggregation?</span></span></p><p cid="n1086" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着人工智能的进步和各行业数字化程度的不断提高，个人数据收集和分析的规模持续增长，导致对个人数据和身份隐私保护的需求日益增加。然而，现有的安全聚合方法（如ACORN (USENIX 2023)）在确保输入数据的隐私和合规性方面表现良好，却无法满足客户端匿名性的要求。简单地应用匿名凭证允许先前已识别的恶意客户端（例如使用不合规数据的客户端）通过更新其凭证重新进入聚合轮次，从而逃避责任。为解决这一问题，我们提出了WhiteCloak，这是首个在客户端匿名性下确保责任归属的安全聚合解决方案。WhiteCloak要求每个客户端i使用匿名凭证$\tilde{i}_{\tau}$参与第$\tau$轮。参与前，每个客户端必须提交一个零知识证明，验证自己未被列入黑名单，从而防止恶意客户端通过更改凭证逃避责任。WhiteCloak可以无缝集成到现有框架中。在SHAKESPEARE数据集的联邦学习实验中，WhiteCloak仅增加了1.77秒的额外处理时间和35.68KB的通信开销，分别占ACORN总开销的0.34%和0.1%。</span></span></p><p cid="n1087" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-s142-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-s142-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1089" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">263、WiFinger: Fingerprinting Noisy IoT Event Traffic Using Packet-level Sequence Matching</span></span></p><p cid="n1090" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">智能家居等物联网环境容易受到隐私推断攻击，攻击者可以通过分析加密网络流量模式来推断设备状态甚至人的活动。虽然大多数现有攻击利用机器学习技术来发现此类流量模式，但由于无线流量（尤其是Wi-Fi）的高噪声性和数据包丢失问题，它们在无线流量上的表现不佳。此外，这些方法通常针对区分分块的物联网事件流量样本，无法有效同时跟踪多个事件。在这项工作中，我们提出了WiFinger，一种针对噪声流量的细粒度多物联网事件指纹识别方法。WiFinger将流量模式分类任务转化为子序列匹配问题，并引入了新技术来处理高时间复杂度，同时保持高准确性。此外，它对训练样本量的依赖减少了未来指纹更新的工作量。实验表明，在实际威胁模型下，WiFinger的性能优于现有方法，平均召回率达到89%（相比之下，分别为49%和46%），且对于各种物联网事件的误报率几乎为零。</span></span></p><p cid="n1091" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f1083-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f1083-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1093" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">264、XR Devices Send WiFi Packets When They Should Not: Cross-Building Keylogging Attacks via Non-Cooperative Wireless Sensing</span></span></p><p cid="n1094" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">随着扩展现实(XR)技术不断融入各个领域，各种安全漏洞——如按键推断(键盘记录)——已成为日益增长的担忧。几种键盘记录攻击证明了利用语音和视觉等多种模态利用此漏洞的可行性。然而，这些攻击通常需要视线(LoS)和/或近距离(&lt;10米)的限制。我们提出了一种针对XR设备的新型键盘记录攻击，利用WiFi无线传感。与先前方法不同，我们的攻击不需要视线，并且在各种场景中均有效，包括远距离、跨建筑物设置(最远30米)。我们的攻击仅需一个廉价、口袋大小的接收装置即可收集受害者的WiFi数据包。与利用WiFi的先前键盘记录攻击相比，我们的方法首次消除了对独立发射器和接收器或虚假热点的需求。因此，与先前方法不同，我们的攻击即使在远距离也有效。核心思想在于利用WiFi芯片组中的安全漏洞。此漏洞允许攻击者向受害者设备发送一个虚假的未加密数据包，作为响应，受害者设备会不由自主地自动传输一个确认(ACK)数据包。通过利用此机制，我们可以持续强制头显的WiFi芯片组传输数据包，从而从受害者的头显中收集大量信道状态信息(CSI)数据。随后，我们开发了一种新颖的无监督信号处理算法，利用CSI数据进行姿态估计，定位受害者的手和手指，最终实现按键推断。我们在Meta Quest 2和Meta Quest 3头显上评估了我们的攻击，测试条件多样，包括距离从1米到30米，角度从-90°到+90°，多用户场景，以及穿墙场景，证明了其在广泛环境中的鲁棒性和有效性。我们的攻击在建筑物内实现了78.6%的top-25准确率，可推断长达15个字符的密码。</span></span></p><p cid="n1095" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f926-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f926-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p cid="n1097" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">265、ZKSL: Verifiable and Efficient Split Federated Learning via Asynchronous Zero-Knowledge Proofs</span></span></p><p cid="n1098" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">在垂直联邦学习(VFL)中，先前的工作主要侧重于保护数据隐私，而忽视了参与者可能操纵本地模型执行以实施完整性攻击的风险。将零知识证明(ZKPs)整合到训练过程中可以确保各方计算的可验证性，同时不泄露私有数据。然而，由于以下原因，将深度模型训练直接编码为整体ZKP电路是不切实际的：(i)复杂的电路设计和频繁参数承诺带来的高开销，(ii)嵌入层(跨方信息接口)的证明生成成本高昂，以及(iii)同步证明生成会阻塞迭代训练轮次。为应对这些挑战，我们提出了ZKSL，这是一个高效且异步的VFL框架，在恶意威胁模型下实现可验证训练。ZKSL将深度神经网络划分为分层电路并并行生成其证明，通过&#34;隐私承诺PLONK&#34;(PC-PLONK)确保输入-输出一致性，这是一种轻量级扩展，支持低成本、逐次迭代的参数承诺。对于嵌入层，ZKSL采用概率验证技术，将证明复杂度从${O(Nnd)}$降低到${O(nd)}$。此外，ZKSL集成了异步计算-证明调度机制，将证明生成与训练迭代解耦，有效缓解了流水线停滞问题。在DeepFM和CNN模型上的实验结果表明，ZKSL将证明生成时间最多减少73%，同时保持99.4%的准确率，展示了其在实际联邦学习中的卓越可扩展性和实用性。</span></span></p><p cid="n1099" mdtype="paragraph" style="box-sizing: border-box;"><span md-inline="plain" style="box-sizing: border-box;"><span leaf="">下载链接：</span></span><span md-inline="url" spellcheck="false" style="box-sizing: border-box;"><span leaf=""><a href="https://www.ndss-symposium.org/wp-content/uploads/2026-f2008-paper.pdf" target="_blank">https://www.ndss-symposium.org/wp-content/uploads/2026-f2008-paper.pdf</a></span></span></p><hr style="box-sizing: content-box;"/><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=5ac174a8&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486060%26idx%3D3%26sn%3Db2d8e7d8dbdb8671d797ec81df76296f">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sun, 01 Mar 2026 14:04:00 +0800</pubDate>
    </item>
    <item>
      <title>2025 年终推荐书单</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486048&amp;idx=1&amp;sn=a97e97b359138ee2bdea8c69bb6c2e10</link>
      <description>今年共读完 37 本书，其中文学类的书籍占比较多，选择个人最喜欢的 16 本书做推荐，刚好推荐的书都在微信读书上。&#xA;&#xA;这次就不写书评了，附上的图片上面都有二维码，直接点开就可以看到一些推荐值和书评了。&#xA;&#xA;如果要对 2025 年的读书做一个总结的话，那主要有以下几个感受：&#xA;&#xA;1.随着年龄的增长与生活阅历的丰富，自己越来越喜欢阅读文学作品，因为它能够抚慰人心，能够表达你表达不了的话，犹如哑巴突然开口说话一般惊喜；其次，它能够给你丰富的想象，对从事理工科研究工作的人，可以培养创新思维。&#xA;&#xA;2.网络安全书籍读的越来越少，因为很多同类书籍大都出版过，已经鲜有新主题方向的书可出版，导致此类新书就越来越少了；其次，个人认为网络安全研究越往深处走，就会回归到底层的一些计算机基础上，所以近两年看的计算机专业基础书籍反而更多；最后是时效性问题，书籍的出版对于内容很多时候它容易过时，而论文的出版则更为及时，所以现在我读论文的数量反而更多。&#xA;&#xA;3.给自己创造读书的空间很重要，比如买几个阅读器和书架，放在不同的地方，挑把坐的舒服的椅子，便于自己随手随地可以看书，同时可以培养孩子的阅读习惯。</description>
      <content:encoded><![CDATA[<p><span>漏洞战争</span> <span></span> <span style="display: inline-block;">广东</span></p>






  
  
  <p>今年共读完 37 本书，其中文学类的书籍占比较多，选择个人最喜欢的 16 本书做推荐，刚好推荐的书都在微信读书上。</p><p>这次就不写书评了，附上的图片上面都有二维码，直接点开就可以看到一些推荐值和书评了。</p><p>如果要对 2025 年的读书做一个总结的话，那主要有以下几个感受：</p><p>1.随着年龄的增长与生活阅历的丰富，自己越来越喜欢阅读文学作品，因为它能够抚慰人心，能够表达你表达不了的话，犹如哑巴突然开口说话一般惊喜；其次，它能够给你丰富的想象，对从事理工科研究工作的人，可以培养创新思维。</p><p>2.网络安全书籍读的越来越少，因为很多同类书籍大都出版过，已经鲜有新主题方向的书可出版，导致此类新书就越来越少了；其次，个人认为网络安全研究越往深处走，就会回归到底层的一些计算机基础上，所以近两年看的计算机专业基础书籍反而更多；最后是时效性问题，书籍的出版对于内容很多时候它容易过时，而论文的出版则更为及时，所以现在我读论文的数量反而更多。</p><p>3.给自己创造读书的空间很重要，比如买几个阅读器和书架，放在不同的地方，挑把坐的舒服的椅子，便于自己随手随地可以看书，同时可以培养孩子的阅读习惯。</p>
  <p><img src="https://wechat2rss.xlab.app/img-proxy/?k=f11b6082&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1STh2ZU6wXbQZHeOiaQnCXVFFvA82dWDCJ7AXj9HbohicIRCnkBjGlnXibQ%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=9a2df137&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SWhHa7JiaL0EOmMy0nsjsHjCyzKQUDYzC1yeh1nDcSu0DibKczSpzdffg%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=32238130&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SFP71LVZEQt8HmKQGwL7GIMRqj9SWZTmu5ZQbqPcAWDgnibaUAZHSWicw%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=cf8413c2&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1Sn23PD3Z0OeoHR03hKYmabVTWn9WuON5tGnaV1j1N4Q0X83Kuic5gumQ%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=8365735b&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SLJqoV5so4TsiblDTHBRdg3XYH06lib5kACNhac0AN2sl2lhWqPaFpz4w%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=505a1741&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SHFn2NLSb7dMBiaG7HT0Q7AIg2gXbexiabKdcic33MLvI397MWG8QPet5g%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=07679c79&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SYBjPJwN2eD6aGX3f3RkTActfLezLIGhUADuNib5jh1WcW6RXkdibj5AQ%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=1a74d68a&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SyjB8oE80oB19z3ia2nIu2B38NiaTZricpBwNOclKXhB2JYic5fciaQUlOWg%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=a9bca9eb&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SZb3zToenbibibcic4A2ZtiaOYRqEPcWEXrsicibh7RgfRq7PfZbdUn5SlMfQ%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=e3c2caed&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SjFtIYtfibR6tBUFibibmAtHvowfiaUhPWWU5K9g8uFmsIzHS9FXwo1Bmhg%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=13392c30&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SG0Ldb5jn3lZkobhnuibwnDVr7HWlDKtDGpjEnkhh8aiceXOty8TicXSqQ%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=2671b9ed&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SicpMjQHxkIDibUZ47NTAzBvRqUfUIDN5KFWX9lLjxmrIAzVo01RGicZuw%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=2d9147a5&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SZ58ye8Iu2kHkMRTicTbMfrUpibI6ibcG2Qy9LV8QBOT1FASStVmNfibdEw%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=def5fa95&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SF1BnrxdzYC1H5KkMb53hVplVdLoKeFiaZTic3ngS8BULr5pKp2l9hzjA%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=67d8c1a7&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1SBal0Exyew7DJOK5W91QuicZKr5ibfqPk8RpmK2MvcOoZ8uaicalP4RN9w%2F0%3Fwx_fmt%3Djpeg"/></p><p><img src="https://wechat2rss.xlab.app/img-proxy/?k=17d771db&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdUoS67lwDq2IXv9EFdF8P1S014G9LcOq37bCuDrPIIljfyImpO91ZjFP6U6DicicvVXrqSWL4PTCQJw%2F0%3Fwx_fmt%3Djpeg"/></p>



<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=5bdf18b3&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486048%26idx%3D1%26sn%3Da97e97b359138ee2bdea8c69bb6c2e10">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Fri, 02 Jan 2026 08:54:04 +0800</pubDate>
    </item>
    <item>
      <title>软件测试顶会ISSTA 2025 论文清单与摘要（补遗）</title>
      <link>https://mp.weixin.qq.com/s?__biz=MzU0MzgzNTU0Mw==&amp;mid=2247486003&amp;idx=1&amp;sn=82d1280ff69952f09d94eb5f9ff2d59a</link>
      <description>ISSTA 主会场 Full Paper，之前发的只是附属研讨会的短文</description>
      <content:encoded><![CDATA[<p>
<span>漏洞战争</span> <span>2025-09-20 21:58</span> <span style="display: inline-block;">广东</span>
</p>

<p>ISSTA 主会场 Full Paper，之前发的只是附属研讨会的短文</p>



<p>
<img src="https://wechat2rss.xlab.app/img-proxy/?k=e96d43b2&amp;u=https%3A%2F%2Fmmbiz.qpic.cn%2Fmmbiz_jpg%2FicNlicgdbzSdVHDOVlQ0Jn7HjKFFqODGbv5zzCpzBVDT1T4I2vHicCKRaewJut6yNUwaOFST1VIKERCxYDsXV2wGg%2F0%3Fwx_fmt%3Djpeg"/>
</p>


<p class="mp_profile_iframe_wrp" nodeleaf=""><mp-common-profile class="js_uneditable custom_select_card mp_profile_iframe" data-pluginname="mpprofile" data-nickname="漏洞战争" data-alias="vulwar" data-from="0" data-headimg="http://mmbiz.qpic.cn/mmbiz_png/icNlicgdbzSdWzbtNBGKasvuCIJ0vjJMt3QXRbMdakfbN6oq553ax43vZeJaD0QPnP4ktdfDS01vozNKsiapNz0SQ/0?wx_fmt=png" data-signature="谈人生，聊梦想，话安全，说风云" data-id="MzU0MzgzNTU0Mw==" data-is_biz_ban="0" data-service_type="1" data-verify_status="1"></mp-common-profile></p><h3 cid="n0" mdtype="heading" data-pm-slice="0 0 []"><span leaf="" style="color:rgba(0, 0, 0, 0.9);font-size:17px;font-family:&#34;mp-quote&#34;, &#34;PingFang SC&#34;, system-ui, -apple-system, BlinkMacSystemFont, &#34;Helvetica Neue&#34;, &#34;Hiragino Sans GB&#34;, &#34;Microsoft YaHei UI&#34;, &#34;Microsoft YaHei&#34;, Arial, sans-serif;line-height:1.6;letter-spacing:0.034em;font-style:normal;font-weight:normal;">前一篇发的其实是</span><span leaf="">ISSTA Companion 短文，之前没注意，下载好多篇看了之后，发现都是短文，质量远不及主会场的论文，今天重新采集主会场 Full Paper 列表发出来。</span></h3><h3 cid="n0" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; " data-pm-slice="0 0 []"><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">1、A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit Testing</span></span></h3><p cid="n2" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">单元测试在软件开发中扮演着关键角色，能有效提升软件质量与可靠性。然而人工生成高效测试用例耗时费力，这推动了单元测试自动化研究的发展。近年来，大语言模型（LLMs）在测试生成、断言生成和测试演进等单元测试任务中展现出潜力，但现有研究范围有限且缺乏对LLMs效能的系统评估。为填补这一空白，我们开展了针对单元测试任务的大规模大语言模型微调实证研究。本研究涵盖三项单元测试任务、五个基准数据集、八项评估指标以及37种不同架构和规模的流行LLMs，累计消耗超过3,000个NVIDIA A100 GPU小时。我们聚焦三个核心研究问题：（1）LLMs相较于现有最优方法的性能表现；（2）不同因素对LLM性能的影响；（3）微调与提示工程的效果对比。研究发现：在所有三项单元测试任务中，LLMs在几乎全部指标上均优于现有最优方法，凸显了微调LLMs在单元测试任务中的潜力。进一步地，大规模仅解码器模型在所有任务中表现最佳，而编码器-解码器模型在相同参数规模下性能更优。此外，通过对比微调与提示工程的性能表现，我们发现提示工程方法在单元测试任务中具有显著潜力。我们继而探讨了测试生成任务中的关键问题，包括数据泄露问题、缺陷检测能力和指标对比。最后，我们进一步为近期基于LLM的单元测试任务实践提出了具体指导准则。总体而言，本研究证明了微调LLMs在单元测试任务中的广阔前景，并有效降低了实际场景中单元测试专家的人工投入成本。</span></span></p><p cid="n3" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728951" target="_blank">https://doi.org/10.1145/3728951</a></span></span></p><h3 cid="n4" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">2、A Low-Cost Feature Interaction Fault Localization Approach for Software Product Lines</span></span></h3><p cid="n5" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">在软件产品线（SPL）中，定位缺陷特征交互能帮助开发人员识别测试失败的根源，从而减轻其工作负担。由于潜在交互数量随特征数量呈指数级增长，该任务面临巨大挑战——尤其对于大型SPL而言，搜索空间极为庞大。现有方法通过基于可疑特征选择（例如出现在失败配置但未通过测试的特征）构建和检测潜在特征交互，部分解决了这一问题。然而这些方法往往忽略缺陷特征交互与测试失败之间的因果关系，导致搜索空间过大和故障定位成本高昂。为此，我们提出一种基于反事实推理的低成本故障定位方法（CRFL），通过缩减搜索空间和减少冗余计算来提升定位效率。具体而言，CRFL运用反事实推理推断可疑特征选择，并采用对称不确定性过滤无关特征交互。此外，该方法融合两项发现机制以避免重复生成和检测相同特征交互。我们在八个公开SPL系统上评估本方法性能，并针对BerkeleyDB和TankWar生成多个缺陷变异体以支持大规模真实SPL的对比实验。实验结果表明：对于小型SPL（6-9个特征），本方法将搜索空间缩减51%∼73%；对于大型SPL（13-99个特征），缩减幅度达71%∼88%。本方法平均运行时间比现有最优技术快约15.6倍。当与语句级定位技术结合时，CRFL能高效定位缺陷语句，这证明其可准确识别缺陷特征交互。</span></span></p><p cid="n6" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728917" target="_blank">https://doi.org/10.1145/3728917</a></span></span></p><h3 cid="n7" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">3、ALMOND: Learning an Assembly Language Model for 0-Shot Code Obfuscation Detection</span></span></h3><p cid="n8" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">代码混淆是一种通过增加软件理解和逆向工程难度来保护软件的技术。然而，该技术也可能被恶意利用，如实施代码抄袭或开发恶意程序。基于学习的技术在监督学习和标注训练集的帮助下已取得显著成功。但面对现实环境中涉及私有开发且未公开的混淆器时，这些监督学习方法在面对未见未知类别的混淆技术时，其泛化性和鲁棒性常引发担忧。</span></span></p><p cid="n9" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出ALMOND——一种用于检测二进制可执行文件中代码混淆的新型零样本方法。与先前监督学习方法不同，ALMOND无需标注的混淆样本进行训练，而是利用仅在未混淆汇编代码上预训练的语言模型来识别混淆引入的语言偏差。其核心创新是采用&#34;错误困惑度&#34;作为检测指标，该指标专注于模型未能预测的标记。连续错误困惑度进一步强化此方法，以捕捉混淆序列特有的连续预测错误特征。</span></span></p><p cid="n10" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">实验表明，ALMOND对未见混淆方法的检测准确率达96.3%，优于监督基线方法。在真实恶意软件样本上，其AUC值达到0.869，显著超越监督学习基线。我们的数据集、预训练模型及评估代码将在</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://github.com/palmtreemodel/ALMOND" target="_blank">https://github.com/palmtreemodel/ALMOND</a></span></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""> 公开。</span></span></p><p cid="n11" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728886" target="_blank">https://doi.org/10.1145/3728886</a></span></span></p><h3 cid="n12" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">4、Adding Spatial Memory Safety to EDK II through Checked C (Experience Paper)</span></span></h3><p cid="n13" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">嵌入式软件主要采用C语言编写，由于空间内存问题而易受内存破坏漏洞影响。尽管存在多种内存安全技术，但由于资源限制和缺乏标准化操作系统支持，这些技术通常不适用于嵌入式系统。Checked C作为一种向后兼容的内存安全C方言，通过使用指针注解进行运行时检查，以最小开销提升空间内存安全性，提供了潜在解决方案。本文首次呈现了将典型嵌入式代码库EDK2（开源UEFI实现）移植到Checked C的实践报告，重点阐明移植过程中的挑战，并为在类似嵌入式系统中应用Checked C提供见解。我们还开发了增强型自动注解工具e3c，将转换率提升25%，显著简化了向Checked C的转换过程。</span></span></p><p cid="n14" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728929" target="_blank">https://doi.org/10.1145/3728929</a></span></span></p><h3 cid="n15" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">5、AdverIntent-Agent: Adversarial Reasoning for Repair Based on Inferred Program Intent</span></span></h3><p cid="n16" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">自动程序修复（APR）技术已展现出显著成果，尤其是神经网络的应用。当前多数APR工具聚焦于测试套件规定的代码转换，而非对程序意图和高级错误规约的推理。若缺乏对程序意图的准确理解，这些工具易生成过度拟合不完整测试套件的补丁，无法体现开发者真实意图。然而，程序意图推理本身极具挑战性。本研究提出一种基于批判与对抗推理的方法——AdverIntent-Agent。其创新性在于将重心从生成多个APR补丁转向推断多种潜在程序意图。理想情况下，我们致力于推断出具有一定对抗性的多重意图，从而最大化至少一种意图与开发者原始意图高度匹配的概率。AdverIntent-Agent采用多智能体架构，包含推理智能体、测试智能体和修复智能体：推理智能体首先生成对抗性程序意图及对应错误语句；测试智能体随后为每个推断意图生成对抗测试用例，构建使用相同输入但预期输出不同的测试预言；最终修复智能体通过动态精准的LLM提示生成同时满足推断程序意图和生成测试的补丁。我们在Defects4J 2.0和HumanEval-Java基准上评估AdverIntent-Agent，分别成功修复77和105个错误。本研究通过让开发者以自然语言评估程序意图（而非审查代码补丁），显著降低了补丁审核所需的工作量。</span></span></p><p cid="n17" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728939" target="_blank">https://doi.org/10.1145/3728939</a></span></span></p><h3 cid="n18" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">6、An Investigation on Numerical Bugs in GPU Programs Towards Automated Bug Detection</span></span></h3><p cid="n19" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">通用图形处理器（GPU）计算已成为主流的并行计算范式，在科学计算和深度学习等多个领域带来显著的性能提升。然而，GPU程序易受数值错误影响，可能导致计算结果错误或系统崩溃。这类错误的检测、调试和修复极具挑战性：它们依赖于特定输入值或类型，缺乏可靠的错误检查机制和验证基准，且GPU独特的编程规范增加了定位根本原因的难度。修复过程还需要掌握GPU计算及数值库的领域知识。因此，深入理解GPU数值错误（GPU-NBs）的特征对开发有效解决方案至关重要。</span></span></p><p cid="n20" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文通过分析GitHub中397个真实错误样本，对GPU-NBs展开全面研究。我们归纳了常见根本原因、症状、触发错误的输入模式与测试验证方法，并总结了修复策略。同时，我们开发了初步检测工具GPU-NBDetect，可检测六类不同数值错误。该工具在四个数值库的186个数学函数中共计发现226个错误，其中60个已获开发者确认。本研究为GPU数值错误的检测与预防技术奠定了基础，并为构建高效调试与自动修复工具提供了重要参考。</span></span></p><p cid="n21" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728950" target="_blank">https://doi.org/10.1145/3728950</a></span></span></p><h3 cid="n22" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">7、Are Autonomous Web Agents Good Testers?</span></span></h3><p cid="n23" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">尽管自动化测试技术不断进步，但由于测试脚本脆弱性带来的高维护需求——应用程序结构的微小变更就可能导致脚本失效，手动测试仍然占据主导地位。大型语言模型（LLMs）的最新发展为自主网页代理（AWAs）提供了潜在替代方案，这些代理能够自主与应用程序进行交互。此类代理可作为自主测试代理（ATAs），通过使用类似人类测试人员所需的自然语言指令，有望减少对高维护性自动化脚本的依赖。本文研究了将AWAs应用于自然语言测试用例执行的可行性及其评估方法。</span></span></p><p cid="n24" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们的贡献包括：（1）构建包含三个离线Web应用程序的基准测试集及113个手动测试用例（含通过/失败案例），用于评估比较ATAs性能；（2）开发SeeAct-ATA和pinATA两个开源ATA实现，能够执行测试步骤、验证断言并给出判定结果；（3）通过基准测试进行对比实验，量化评估ATA的有效性。最后我们还对性能最佳的PinATA进行了定性评估以识别其局限性。</span></span></p><p cid="n25" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">研究结果表明：在执行测试用例时，我们简单的SeeAct-ATA实现相比更先进的PinATA实现性能较差（性能差距达50%）。尽管PinATA能获得约60%的正确判定率和高达94%的特异性指标，但我们发现要开发更具韧性和可靠性的ATAs仍需解决若干局限性，这为构建强健、低维护的测试自动化系统指明了方向。</span></span></p><p cid="n26" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728879" target="_blank">https://doi.org/10.1145/3728879</a></span></span></p><h3 cid="n27" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">8、AudioTest: Prioritizing Audio Test Cases</span></span></h3><p cid="n28" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于深度神经网络（DNN）的音频分类系统是影响日常生活各类应用（如语音助手）的核心组件。确保此类系统的准确性至关重要，因为分类错误可能导致严重的安全问题与用户信任危机。然而音频分类器的测试面临重大挑战：音频测试样本的人工标注成本极高。测试优先级排序已成为缓解标注成本问题的有效手段，该方法通过优先处理可能被误分类的测试样本，实现关键样本的早期标注，从而提升调试效率。但现有优先级排序方法在音频测试样本上存在局限：1）基于代码覆盖的方法在效果和效率上均逊于基于置信度的方法；2）基于置信度的方法仅依赖预测概率向量，忽略了音频数据的独特性；3）基于变异的方法缺乏针对音频设计的变异操作，难以适用于音频测试样本。为此，我们提出专为音频测试样本设计的新型优先级排序方法AudioTest。其核心思想是：与误分类样本空间距离越近的测试越可能被误分类。基于音频数据的特性，AudioTest生成四类特征：时域特征、频域特征、感知特征和输出特征。该方法将每项测试的四类特征拼接为特征向量，并采用精心设计的特征变换策略，使误分类样本在空间中的分布更紧凑。AudioTest借助训练好的模型，根据变换后的向量预测每项测试的误分类概率，并据此排序。我们在包含纯净与带噪数据集的96个实验对象上评估AudioTest，采用故障检测率（PFD）和平均故障检测百分比（APFD）两项经典指标。结果表明，AudioTest在PFD和APFD上均优于所有对比方法。在纯净数据集上，相比基线方法的平均提升幅度为12.63%至54.58%；在带噪数据集上，提升幅度为12.71%至40.48%。</span></span></p><p cid="n29" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728907" target="_blank">https://doi.org/10.1145/3728907</a></span></span></p><h3 cid="n30" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">9、Automated Attack Synthesis for Constant Product Market Makers</span></span></h3><p cid="n31" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">去中心化金融（DeFi）实现了传统金融中许多前所未有的新型应用，但同时也引入了新型安全漏洞。此类漏洞的一个典型代表是代币合约与遵循恒定乘积做市商（CPMM）模型的去中心化交易所（DEX）之间的可组合性缺陷。我们将这类缺陷称为CPMM可组合性漏洞，其根源在于代币合约的设计问题导致其与CPMM模型不兼容，进而危及CPMM生态系统中的其他代币。自2022年以来，此类漏洞已引发23次攻击事件，累计造成220万美元损失。智能合约审计公司BlockSec报告显示，仅2023年2月就发生了138次此类攻击。</span></span></p><p cid="n32" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出CPMMX工具，能够自动检测整个区块链上的CPMM可组合性漏洞。为实现这种可扩展性，我们首先形式化定义了CPMM可组合性漏洞，发现破坏两个安全不变量即可诱发此类漏洞。基于该发现，我们设计了采用&#34;先浅层后深度&#34;双步检测机制的CPMMX工具：首先通过浅层搜索识别破坏不变量的交易，继而通过深度搜索精炼这些交易以验证攻击者可获利性。我们在两个公共数据集和一个合成数据集上使用五种基线方法进行评估。实验表明，CPMMX的漏洞检测数量是基线方法的1.5至2.5倍，分析速度显著提升且F1分数更高。此外，我们将CPMMX应用于以太坊和币安网络最新区块的所有合约，新发现26个可获利漏洞，潜在攻击总收益达1.57万美元。</span></span></p><p cid="n33" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728872" target="_blank">https://doi.org/10.1145/3728872</a></span></span></p><h3 cid="n34" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">10、Automated Test Transfer across Android Apps using Large Language Models</span></span></h3><p cid="n35" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">移动应用在日常生活中的普及要求采用强大的测试策略来确保质量和效率，特别是通过基于使用的端到端移动应用用户界面（UI）测试。然而，手动创建和维护此类测试对开发者而言成本高昂。由于许多应用在多样化UI下具有相似功能，先前研究已证明在同领域不同应用间迁移UI测试的可能性，从而避免了手动编写测试的需求。但这些方法难以适应现实场景中的变化，当源应用和目标应用相似度不高或未能准确迁移测试预言时往往存在局限。本文提出创新技术LLMigrate，利用大语言模型（LLM）高效实现跨移动应用的基于使用的UI测试迁移。实验评估表明，LLMigrate在自动化测试迁移中可实现97.5%的成功率，将手动编写测试的工作量减少91.1%。相较于现有最佳技术，该方案在成功率上提升9.1%，在工作量减少上提高38.2%，为自动化测试迁移设立了新基准。</span></span></p><p cid="n36" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728975" target="_blank">https://doi.org/10.1145/3728975</a></span></span></p><h3 cid="n37" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">11、Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMs</span></span></h3><p cid="n38" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">尽管加密API误用的自动化检测已取得显著进展，但由于依赖手动定义模式，其在复杂目标上的精确度仍然有限。大语言模型（LLM）凭借其上下文感知能力为弥补这一缺陷提供了新途径，但其随机性及幻觉问题对精准安全分析应用构成挑战。本文首次系统研究LLM在加密API误用检测中的应用，并获得重要发现：直接应用LLM的不稳定性导致初始报告中超过半数误报。然而，通过将检测范围与现实场景对齐并采用创新的代码与分析验证技术，可显著提升基于LLM检测的可靠性，实现近90%的检测召回率，这一提升远超传统方法，并成功在成熟基准测试中发现未知漏洞。研究同时揭示了当前LLM存在的共性失效模式，包括密码学知识缺失和代码语义误判等盲点。基于这些发现，我们部署了LLM检测系统，在开源Java和Python代码库（含Apache等知名项目）中新发现63个漏洞（47个已确认，7个已完成修复）。</span></span></p><p cid="n39" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728875" target="_blank">https://doi.org/10.1145/3728875</a></span></span></p><h3 cid="n40" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">12、BinDSA: Efficient, Precise Binary-Level Pointer Analysis with Context-Sensitive Heap Reconstruction</span></span></h3><p cid="n41" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">指针分析是二进制代码逆向工程领域的基础组件。它可用于重建二进制程序的调用图，并能进一步应用于各种安全分析。然而，二进制代码中符号和类型信息的缺失给有效的指针分析带来了巨大挑战。现有研究在对二进制代码进行指针分析时通常采用近似方法，但这些方法往往效率低下且会产生大量误报目标。本文提出了一种专为二进制指针分析定制的新模型BinDSA，该模型将精确性和效率置于完备性之上。它具备字段敏感性和上下文敏感性，采用基于统一化的技术并重建上下文敏感的堆结构。通过联合恢复数据结构和指向关系，进一步提升了分析精度。评估结果表明，BinDSA的效率比当前最先进技术提升5倍，且在未显著牺牲完备性的情况下显著提高了精确度。我们还将BinDSA应用于CVE可达性分析和漏洞检测，证明了其在安全任务中的有效应用。</span></span></p><p cid="n42" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728928" target="_blank">https://doi.org/10.1145/3728928</a></span></span></p><h3 cid="n43" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">13、BinQuery: A Novel Framework for Natural Language-Based Binary Code Retrieval</span></span></h3><p cid="n44" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">二进制函数检索（BFR）在逆向工程中至关重要，用于识别二进制代码中的特定功能，尤其是与恶意行为或漏洞相关的功能。传统的BFR方法依赖启发式规则，往往缺乏处理大规模或多样化二进制分析任务所需的效率和适应性。为应对这些挑战，我们提出了BinQuery——一个基于自然语言的BFR（NL-based BFR）框架，通过自然语言查询以更高的灵活性和精确度检索相关二进制函数。BinQuery引入了创新技术来弥合二进制代码与自然语言之间的信息鸿沟，实现细粒度对齐以提升检索准确度，并利用大语言模型（LLM）优化查询和生成多样化描述。大量实验表明，BinQuery显著超越当前最先进方法，在可比基准测试中召回率@1提升42.55%，性能提高4倍。</span></span></p><p cid="n45" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728927" target="_blank">https://doi.org/10.1145/3728927</a></span></span></p><h3 cid="n46" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">14、Bridge the Islands: Pointer Analysis for Microservice Systems</span></span></h3><p cid="n47" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">微服务架构通过将应用程序分解为松散耦合的服务，为企业级软件带来了可扩展性与灵活性的革命。然而这种范式转变为指针分析——一种对支持各类客户端分析至关重要的基础静态分析技术——带来了独特挑战。现有基础分析方法主要针对单体式企业应用设计，难以处理复杂的服务间通信（如远程过程调用和基于消息的通信）以及依赖注入和Web端点配置等核心编程范式。本文提出Micans，这是首个专门针对微服务系统这些挑战设计的指针分析方案，能够构建跨服务的完整值流。我们在多个领域的真实基准测试上对Micans进行了全面评估，重点关注其在解析服务通信、构建调用图等关键程序信息以及支持污点分析等客户端分析方面的有效性。Micans持续显著优于现有最先进方法，证明了其处理复杂跨服务通信和多样化编程范式的能力。这些结果凸显了Micans作为强大基础分析方案的潜力，推动了静态分析能力以适应现代微服务复杂性的发展。</span></span></p><p cid="n48" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728896" target="_blank">https://doi.org/10.1145/3728896</a></span></span></p><h3 cid="n49" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">15、Bridging the Gaps between Graph Neural Networks and Data-Flow Analysis: The Closer, the Better</span></span></h3><p cid="n50" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">近年来，深度神经网络在编程任务中的应用取得了显著实践成果，这促使研究者开始探索这些模型执行传统程序分析技术的能力。数据流分析（DFA）作为经典且成熟的分析方法，为评估神经网络在此领域的能力提供了契机。基于DFA与图神经网络（GNN）的结构相似性，我们深入探究GNN在多大程度上能有效建模DFA算法。依托神经算法推理（NAR）中的算法对齐概念，我们识别出两大关键挑战：DFA中位向量的非干扰特性，以及算法不同阶段外部信息的复杂处理机制。针对这些不足，我们提出三种逐步与DFA算法对齐的GNN架构——DFA-GNN−、DFA-GNN和DFA-GNN+。实验评估重点关注模型的泛化能力，特别是在小规模样本训练、大规模输入测试场景下的表现。结果表明，具有更高算法对齐度的GNN（如DFA-GNN+）展现出卓越的泛化能力和样本效率，仅需极少训练数据即可精准处理10倍规模的输入。值得注意的是，仅通过输入-输出对训练的GNN模型，其性能可与采用完整执行轨迹监督（当前NAR研究常用方法）的模型相媲美。这一发现凸显了当GNN与目标算法实现算法对齐时，在推理任务中表现出的高效性与鲁棒性。</span></span></p><p cid="n51" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728906" target="_blank">https://doi.org/10.1145/3728906</a></span></span></p><h3 cid="n52" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">16、Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering</span></span></h3><p cid="n53" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">近年来，大型语言模型（LLM）已被应用于代码生成等多种软件工程（SE）任务，显著推进了软件工程任务的自动化进程。然而，评估这些由LLM生成的代码与文本质量仍具挑战性。当前广泛使用的Pass@k指标不仅需要大量单元测试和配置环境，导致人力成本高昂，且不适用于LLM生成文本的评估。而BLEU等传统指标仅衡量词汇而非语义相似性，也已受到质疑。为此，学界新兴趋势是采用LLM进行自动化评估，即&#34;LLM即评判员&#34;方法。这类方法被认为能比传统指标更好地模拟人类评估，且无需依赖高质量参考答案。但它们在软件工程任务中与人类评估的真实契合度尚未得到验证。</span></span></p><p cid="n54" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文通过实证研究探讨了用于评估软件工程任务的LLM即评判员方法，重点关注其与人类判断的一致性。我们选取了七种基于通用LLM的评判方法，以及两种专为评估任务微调的LLM。通过在代码翻译、代码生成和代码摘要这三个最新软件工程数据集上生成LLM响应并进行人工评分后，我们引导这些方法对每个响应进行评估。最终将自动评分结果与人工评估进行对比。研究表明：在代码翻译和代码生成任务中，基于输出的评判方法与人类评分的皮尔逊相关系数分别达到81.32和68.51，接近人类评估水平，显著优于传统最佳指标ChrF++的34.23和64.92。此类方法直接引导LLM输出判断结果，且呈现出更接近人类评分模式的平衡分布特征。最后我们提出洞见与启示，指出当前最先进的LLM即评判员方法在某些软件工程任务中具有替代人类评估的潜力。</span></span></p><p cid="n55" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728963" target="_blank">https://doi.org/10.1145/3728963</a></span></span></p><h3 cid="n56" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">17、Causality-Aided Evaluation and Explanation of Large Language Model-Based Code Generation</span></span></h3><p cid="n57" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">尽管代码生成已广泛应用于各种软件开发场景，但生成代码的质量仍无法得到保证。这在基于大语言模型（LLM）的代码生成时代尤为令人担忧——LLMs被视为复杂而强大的黑盒模型，通过高级自然语言规范（即提示词）来生成代码。然而，鉴于LLMs的复杂性和缺乏透明性，有效评估和解释其代码生成能力存在固有挑战。</span></span></p><p cid="n58" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">受因果分析及其软件工程应用领域最新进展的启发，本文提出一种因果驱动的方法来系统分析提示词与代码间的因果关系。该研究面临三个关键技术挑战：(1) 以规范形式表示文本提示词和代码；(2) 建立高层概念与代码特征间的因果关系；(3) 系统分析多样化的提示词变体。针对这些挑战，我们首先提出基于因果图的新型表示方法，对输入提示词中细粒度、人类可理解的概念进行建模。随后利用构建的因果图识别提示词与衍生代码间的因果关系。</span></span></p><p cid="n59" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们通过对四个主流LLMs模型应用12种以上提示词调整策略进行研究，展示了该框架的洞察能力。研究结果表明：我们的技术具有揭示LLM有效性机理、帮助终端用户理解预测结果的潜力。此外，实验证明该方法可通过合理校准提示词，为提升LLM生成代码质量提供可操作的改进建议。</span></span></p><p cid="n60" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728938" target="_blank">https://doi.org/10.1145/3728938</a></span></span></p><h3 cid="n61" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">18、ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation</span></span></h3><p cid="n62" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">摘要近年来，大型语言模型（LLM）显著提升了自动化代码翻译的性能，在多项传统基准测试中其计算准确率可达80%以上。然而，这些基准中的代码样本多为简短、独立、语句/方法级别且算法导向的类型，与实际编程任务存在偏差。因此，LLM在处理日常开发中所编写代码的实际翻译能力仍不明确。</span></span></p><p cid="n63" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为此，我们构建了类级别代码翻译基准ClassEval-T，首次系统评估了当前主流LLM在类级别代码翻译任务上的表现。该基准扩展自知名类级别Python代码生成基准ClassEval，涵盖数据库操作、游戏设计等实际编程主题，并包含字段、方法、库依赖等多样化上下文关联。我们耗费360人时完成了向Java和C++的手动迁移，提供完整代码样本及对应测试套件。随后，我们设计了整体翻译、最小依赖翻译和独立翻译三种类级别代码翻译策略，在ClassEval-T上评估了涵盖商业型、通用型和代码专用型的八个最新LLM（不同系列与参数量）。实验结果表明：与最广泛研究方法级别代码翻译的基准相比，LLM性能出现显著下降；不同模型间表现差异明显，证实ClassEval-T能有效衡量当前LLM能力。我们进一步探讨了不同翻译策略的适用场景，以及LLM在处理类样本时对依赖关系的感知能力。最后，本文对最佳性能LLM产生的1,243个失败案例进行了全面归因分析，为实践指导和未来研究提供启示。</span></span></p><p cid="n64" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728940" target="_blank">https://doi.org/10.1145/3728940</a></span></span></p><h3 cid="n65" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">19、Clause2Inv: A Generate-Combine-Check Framework for Loop Invariant Inference</span></span></h3><p cid="n66" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">循环不变式推断是程序验证中一个基础且具有挑战性的问题。近期研究采用&#34;生成-验证&#34;框架，通过迭代方式在生成步骤产生候选循环不变式，并在验证步骤进行确认。该框架的主要挑战在于每次迭代中产生高质量的候选不变式，以加速推断过程的收敛。我们通过实证发现：由于逻辑连接词的复杂性，现有方法可能难以直接生成完整不变式，但正确循环不变式的所有子句通常已出现在历史生成结果中。这一发现促使我们改进现有框架，提出了&#34;生成-组合-验证&#34;新框架，将循环不变式推断任务分解为子句生成和子句组合两个阶段。具体而言，我们基于新框架提出了一种新型循环不变式推断方法：采用基于大语言模型的子句生成器与反例驱动的子句组合器。子句生成器利用大语言模型生成大量子句；子句组合器则基于历史反例将生成子句组合成不变式。实验表明，该方法显著优于现有循环不变式推断方案：在线性不变式推断任务中解决316题中的312题，在非线性任务中解决50题中的44题，分别比现有基线方法多解决至少93题和16题。该框架具有良好扩展性，可通过将候选不变式拆分为子句的方式，灵活适配当前基于&#34;生成-验证&#34;框架的各类现有方法。评估显示经轻微适配后，我们的方法能同时提升现有方案的效果与效率：例如Code2Inv原本解决210个线性问题（平均耗时137.6秒），改进后可解决252个问题（平均耗时降至17.8秒）。</span></span></p><p cid="n67" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728920" target="_blank">https://doi.org/10.1145/3728920</a></span></span></p><h3 cid="n68" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">20、ConTested: Consistency-Aided Tested Code Generation with LLM</span></span></h3><p cid="n69" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">近年来，大语言模型（LLM）在代码生成领域取得显著进展，能够根据自然语言需求自动生成代码片段。尽管已达到最先进水平，但LLM生成的代码往往存在准确性与可靠性问题，开发者需耗费大量精力进行调试和评估。现有研究提出利用一致性原则：通过选择能通过更多测试（内部一致性）且在多轮生成中表现稳定（外部一致性）的代码。但由于测试用例同样由LLM生成，依赖错误测试的多数投票会导致不可靠结果。为此，我们提出一种轻量级交互框架，通过融入用户反馈有效引导一致性优化。实验表明，该方法仅需极少人工介入即可显著提升性能。我们在每轮迭代中引入代码与测试的&#34;排序-修正-修复&#34;协同进化机制，通过双向迭代提升二者质量，使代码与测试间的一致性投票更可靠。经大量实验验证，ConTested框架在GPT-3.5和GPT-4o等多个LLM上均表现优异：相对GPT-3.5提升32.9%，相对GPT-4o提升16.97%，较当前最先进的后处理技术MPSC提升11.1%。该改进仅需4轮用户交互，人力成本极低。用户研究进一步证实了ConTested的可行性与成本效益，表明其能在不过度增加负担的前提下有效提升代码生成质量。</span></span></p><p cid="n70" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728902" target="_blank">https://doi.org/10.1145/3728902</a></span></span></p><h3 cid="n71" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">21、Copy-and-Paste? Identifying EVM-Inequivalent Code Smells in Multi-chain Reuse Contracts</span></span></h3><p cid="n72" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">随着以太坊上Solidity合约的发展，越来越多的开发者开始在其他兼容区块链上复用这些合约。然而，开发者可能忽视区块链系统设计之间的差异（如Gas机制和共识协议），导致相同合约在不同区块链上无法实现与以太坊一致的执行效果。这种不一致性揭示了复用合约中的设计缺陷，暴露出阻碍代码可复用性的代码坏味，我们将这种不一致性定义为EVM不等价代码坏味。</span></span></p><p cid="n73" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文首次通过实证研究揭示EVM不等价代码坏味的成因与特征。为确保所识别的坏味真实反映开发者关切，我们收集并分析了1,379份安全审计报告和326篇Stack Overflow帖子，这些资料涉及币安智能链（BSC）和Polygon等EVM兼容区块链上的复用合约。采用开放式卡片分类法，我们定义了六类EVM不等价代码坏味。针对自动化检测需求，我们开发了名为EquivGuard的工具。该工具采用静态污点分析识别关键路径，并通过符号执行验证路径可达性。通过对六大区块链上905,948份合约的分析，我们发现EVM不等价代码坏味普遍存在，平均出现率达17.70%。虽然存在代码坏味的合约未必直接导致财务损失或攻击，但其高出现率及涉及的巨额资产管理规模，凸显了复用这些存在坏味的以太坊合约的潜在威胁。因此，建议开发者摒弃复制-粘贴的编程实践，并在复用以太坊合约前检测EVM不等价代码坏味。</span></span></p><p cid="n74" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728921" target="_blank">https://doi.org/10.1145/3728921</a></span></span></p><h3 cid="n75" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">22、CrossProbe: LLM-Empowered Cross-Project Bug Detection for Deep Learning Frameworks</span></span></h3><p cid="n76" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">深度学习（DL）模型可能给底层DL框架带来可靠性挑战。这些框架易存在缺陷，可能导致崩溃或错误结果，尤其在涉及复杂模型架构和高计算需求时。此类框架缺陷会破坏DL应用，影响用户体验并可能造成经济损失。传统测试DL框架的方法难以适应模型结构的庞大搜索空间、多样化的API以及混合编程与硬件环境的复杂性。尽管近期基于大语言模型（LLM）的技术改进了DL框架模糊测试，但其效果高度依赖于输入提示的质量与多样性，而现有提示通常基于单一框架数据构建。  </span></span></p><p cid="n77" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出一种创新方法，通过利用“镜像问题”（即不同框架中因共通功能存在的类似缺陷）来增强DL框架的测试生成。我们的方法基于以下发现：如PyTorch和TensorFlow等DL框架常因依赖项、开发者错误或边缘情况输入而存在共性缺陷。我们开发了CrossProbe工具，利用LLM从某一框架的现有问题中有效学习，并将获得的知识迁移至另一框架的测试用例生成中，从而发现镜像问题，实现跨框架缺陷检测。为克服框架间功能不兼容与实现差异导致的测试用例生成挑战，我们引入了三个处理流程：对齐、筛选和区分。这些流程通过建立API对数据库、过滤不适用案例及强化跨框架差异，降低迁移错误。实验表明，CrossProbe节省了36.3%的生成迭代次数，且相比现有最先进的基于LLM的测试技术，问题迁移成功率提升25.0%。通过迁移知识，CrossProbe检测到24个独特缺陷，其中19个为先前未知且均需依赖深度学习跨框架知识才能识别。</span></span></p><p cid="n78" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728984" target="_blank">https://doi.org/10.1145/3728984</a></span></span></p><h3 cid="n79" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">23、DataHook: An Efficient and Lightweight System Call Hooking Technique without Instruction Modification</span></span></h3><p cid="n80" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">系统调用是用户空间程序与操作系统内核交互的主要接口。通过挂钩系统调用，可以分析和修改用户空间程序的行为。本文提出DataHook，一种针对32位程序的高效轻量级系统调用挂钩技术。与现有系统调用挂钩技术相比，DataHook通过仅修改少量数据元素而不改变任何程序指令，以极低的挂钩开销实现挂钩。这一独特特性不仅避免了二进制重写带来的多线程冲突，还能支持程序应用更高效的用户空间操作系统子系统。然而现有系统调用挂钩技术难以同时满足这些目标：虽然系统调用用户分发（SUD）和ptrace等技术无需重写进程指令，但会引入显著挂钩开销；而低开销技术通常涉及多字节或多指令的二进制重写，这会带来新的问题。DataHook通过利用32位程序在执行系统调用时的特定行为，巧妙解决了这些问题。简言之，与64位程序不同，32位程序在进行系统调用时使用间接调用指令跳转至执行syscall/sysenter的函数。本文通过操纵间接调用过程中涉及的数据依赖关系来实现系统调用挂钩。这一特性普遍存在于基于glibc的Linux系统上的32位程序中（无论运行于x86或x86-64架构），因此DataHook可部署于这些系统。实验结果表明，DataHook将挂钩开销降至现有技术的1/5.4至1/1429.0。当将DataHook应用于服务器程序使其使用用户空间网络协议栈时，服务器性能提升约4.3倍。在Redis中的应用显示，DataHook仅导致4.0%的性能损失，而其他技术会造成8.0%至94.7%的性能损失。</span></span></p><p cid="n81" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728874" target="_blank">https://doi.org/10.1145/3728874</a></span></span></p><h3 cid="n82" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">24、DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction</span></span></h3><p cid="n83" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">水印技术是一种用于识别数据来源的方法，可帮助防止受保护数据集的滥用。现有代码水印方法借鉴后门研究思想，通过嵌入隐蔽触发器作为水印。尽管这些方法对稀释攻击和后门检测具有高度抵抗力，但其鲁棒性尚未得到充分评估。为填补这一空白，我们提出DeCoMa——一种检测与净化代码数据集水印的双通道方法。针对代码水印隐蔽性带来的高壁垒，DeCoMa利用代码的双通道约束将样本泛化并映射至标准化模板，继而通过识别标准化模板内配对元素的异常关联来提取隐藏水印。最后，DeCoMa通过移除所有含检测水印的样本实现数据净化，从而实现受保护代码的静默占用。我们开展大量实验评估DeCoMa的有效性与效率，涵盖14种代码水印类型和3类代表性智能代码任务（共14种场景）。实验结果表明，DeCoMa在14种水印检测场景中均实现100%的稳定召回率，显著优于基线方法。此外，DeCoMa能有效攻击嵌入率低至0.1%的代码水印，且在净化数据集上训练后仍保持可比模型性能。由于无需模型训练即可完成检测，DeCoMa的效率远超所有基线方法，加速比达31.5至130.9倍。该结果呼吁开发更先进的代码模型水印技术，而DeCoMa可为未来评估提供基线标准。</span></span></p><p cid="n84" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728952" target="_blank">https://doi.org/10.1145/3728952</a></span></span></p><h3 cid="n85" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">25、DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled Code</span></span></h3><p cid="n86" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">反编译器在逆向工程（RE）中被广泛用于将已编译的可执行文件转换为人类可读的伪代码，并支持各种安全分析任务。现有的反编译器（如IDA Pro和Ghidra）侧重于提升反编译代码的可读性而非可重编译性，这限制了进一步的程序化应用——例如基于CodeQL的漏洞分析需要可编译的反编译代码版本。近期基于大语言模型（LLM）的改进方法虽然对人类逆向分析师有帮助，但遗憾的是仍遵循相同路径。</span></span></p><p cid="n87" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文首次探索如何利用现成的大语言模型实现可重编译的反编译——自动将反编译器输出修正为可编译版本。我们首先通过试点研究表明：现有基于规则和基于LLM的方法均难以实现此目标。基于这些发现，我们设计了DecLLM：一种基于迭代LLM修复的循环框架，利用静态重编译和动态运行时反馈作为验证机制，逐步修复反编译器输出。我们在主流C基准测试和真实二进制文件上使用GPT-3.5和GPT-4测试DecLLM，结果表明现成LLM可实现约70%的重编译成功率上限（即100个原本不可重编译的反编译输出中70个变为可重编译）。我们还通过基于CodeQL的漏洞分析验证了可重编译代码的实际应用价值，这种分析无法直接在二进制文件上执行。针对剩余30%的困难案例，我们深入分析其错误类型，为未来面向反编译的LLM设计改进提供见解。</span></span></p><p cid="n88" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728958" target="_blank">https://doi.org/10.1145/3728958</a></span></span></p><h3 cid="n89" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">26、DepState: Detecting Synchronization Failure Bugs in Distributed Database Management Systems</span></span></h3><p cid="n90" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">分布式数据库管理系统（DDBMS）对管理大规模分布式数据至关重要。与单节点数据库不同，DDBMS部署于集群环境，将数据分布至多个节点。其同步过程通过应对数据和集群更新来维持数据一致性。由于同步机制复杂度高，同步缺陷不可避免，可能导致数据不一致、事务错误或集群崩溃，严重损害DDBMS的可用性与可靠性。然而目前针对DDBMS同步过程的测试研究相对匮乏。本文提出DepState框架用于检测同步故障缺陷。该框架通过模拟数据分片和动态集群状态的复杂性，建立跨节点表间依赖关系，并系统性地引入受控的集群状态变化。我们在四大DDBMS系统（MySQL NDB Cluster、MySQL InnoDB Cluster、MariaDB Galera Cluster和TiDB Cluster）上应用DepState，发现25个新缺陷（其中13个已确认）。与最先进工具对比表明：DepState在24小时内多发现14个同步故障缺陷，且在同步相关函数的代码行覆盖率上分别比Jepsen、Mallory、SQLsmith、SQLancer和Mozi高出6.13%-66.51%、5.82%-57.28%、14.12%-83.30%、36.81%-83.88%和43.24%-54.28%。</span></span></p><p cid="n91" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728965" target="_blank">https://doi.org/10.1145/3728965</a></span></span></p><h3 cid="n92" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">27、Detecting Isolation Anomalies in Relational DBMSs</span></span></h3><p cid="n93" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">关系数据库管理系统（DBMS）通过事务确保数据一致性与完整性，并提供多种隔离级别以平衡一致性与性能。然而，关系型DBMS中的隔离异常可能破坏其宣称的隔离级别，导致严重后果（例如错误的查询结果和数据库状态）。现有隔离检查器仅能处理简单的键值式数据模型及相关的read(key)/write(key,value)操作，无法直接支持关系数据模型和复杂SQL操作。</span></span></p><p cid="n94" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出新型黑盒式关系型DBMS隔离检查器IsoRel，可支持关系数据模型与复杂SQL操作。为推断关系型DBMS中事务间的依赖关系，我们首先设计了一种与隔离机制无关的SQL语句插桩方案：通过在每张数据库表中使用两个辅助列，记录每条SQL语句访问的数据行。随后利用SQL语句的记录数据构建关系型事务的依赖图，并根据异常模式识别隔离异常。</span></span></p><p cid="n95" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们在五种广泛使用的关系型DBMS（MySQL、PostgreSQL、MariaDB、CockroachDB和TiDB）及其所有支持的隔离级别上评估IsoRel，共发现48种违反Adya所定义隔离级别的独特隔离异常。</span></span></p><p cid="n96" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728953" target="_blank">https://doi.org/10.1145/3728953</a></span></span></p><h3 cid="n97" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">28、Doctor: Optimizing Container Rebuild Efficiency by Instruction Re-orchestration</span></span></h3><p cid="n98" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">容器化技术彻底改变了软件部署方式，其中Docker凭借其易用性和一致的运行时环境成为行业引领者。随着Docker使用量的增长，优化Dockerfile性能（特别是减少重建时间）已成为维持高效CI/CD流水线的关键。然而，现有优化方法主要针对单次构建，未考虑修改和迭代过程中产生的重复重建成本，限制了长期效率收益。为弥补这一缺陷，我们提出Doctor方法——通过指令重排序提升Dockerfile构建效率，该方法攻克了四大核心挑战：识别指令依赖关系、预测未来修改、确保行为等效性以及管理优化计算复杂度。我们基于Dockerfile语法建立了完整的依赖关系分类体系，并通过历史修改分析优先处理频繁修改的指令。Doctor采用加权拓扑排序算法优化指令顺序，在保持功能性的同时最小化未来重建时间。对2,000个GitHub代码库的实验表明，Doctor成功优化了92.75%的Dockerfile，平均降低26.5%的重建时间，其中12.82%的文件实现超50%的降幅。值得注意的是，86.2%的案例保持了功能相似性。这些发现为Dockerfile管理提供了最佳实践，使开发者能通过科学的优化策略提升Docker效率。</span></span></p><p cid="n99" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://dl.acm.org/doi/10.1145/3728870" target="_blank">https://dl.acm.org/doi/10.1145/3728870</a></span></span></p><h3 cid="n100" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">29、Dynamically Fusing Python HPC Kernels</span></span></h3><p cid="n101" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">近年来，高性能计算领域呈现出两大趋势：一是越来越多地采用Kokkos等性能可移植框架，二是Python等解释型语言的普及。PyKokkos顺应这些趋势，允许开发者使用Python编写性能可移植的内核，显著提升了开发效率。然而，开发者仍面临并行代码组织的问题——将代码拆分为独立内核虽能简化测试与调试，但可能导致性能下降。为让开发者自由组织内核的同时确保性能，我们提出PyFuser：一个用于自动融合性能可移植PyKokkos内核的程序分析框架。该框架动态追踪内核调用，并在应用程序请求计算结果时延迟融合内核。通过提升数据复用率、改进编译器优化效果并减少内核启动开销，PyFuser生成的融合内核可实现加速，且无需修改现有PyKokkos代码。我们还引入了自动化代码转换技术，进一步优化PyFuser生成的融合内核。实验表明，在NVIDIA/AMD GPU及Intel/AMD CPU平台上，PyFuser相比未融合内核平均可实现3.8倍的加速比。</span></span></p><p cid="n102" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728959" target="_blank">https://doi.org/10.1145/3728959</a></span></span></p><h3 cid="n103" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">30、Effective REST APIs Testing with Error Message Analysis</span></span></h3><p cid="n104" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">REST API在现代企业系统构建中至关重要，但对其进行有效测试仍存在挑战，尤其是从规范中推断约束条件存在困难。现有测试方法通常依赖HTTP状态码的反馈来指导输入生成，但忽略了伴随错误消息中的宝贵信息，导致探索API输入空间的效果受限。本文提出EmRest——一种利用错误消息分析来增强REST API有效及异常测试输入生成的黑盒测试方法。针对被测操作，EmRest首先识别其每个输入参数所有可能的值分配策略，随后基于这些策略反复应用组合测试来采样测试输入，并通过统计分析收到的错误消息（400系列状态码）来推断并排除无效的值分配策略组合（即输入空间的约束）。此外，EmRest通过变异最终确定的有效值分配策略来生成异常测试输入，并对收到的错误消息（500系列状态码）进行分类以识别易出故障的操作，为其分配更多测试资源。在16个真实REST API上的实验结果表明，EmRest在50%的API中实现了比现有最优方法更高的操作覆盖率，并检测到226个其他方法未发现的独特错误。</span></span></p><p cid="n105" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728964" target="_blank">https://doi.org/10.1145/3728964</a></span></span></p><h3 cid="n106" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">31、Enhanced Prompting Framework for Code Summarization with Large Language Models</span></span></h3><p cid="n107" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">代码摘要技术对于提升软件开发效率至关重要，它能帮助开发者快速理解并维护软件项目。近期研究利用大型语言模型生成精准代码摘要已展现出优异性能，这主要得益于其强大的生成能力。采用连续提示技术的语言模型能够探索更广阔的问题空间，从而释放更大潜力。然而此类方法也面临特定挑战，尤其是在适配任务特定场景方面——而这正是离散提示的优势所在。此外，编程语言与自然语言之间的本质差异会增加语言模型的理解难度，影响复杂编程场景下摘要的准确性与相关性。这些问题可能导致输出结果与实际需求不匹配，凸显出需要进一步研究以增强语言模型在代码摘要中的有效性。  </span></span></p><p cid="n108" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为突破这些局限，我们融合上述两种方法的优势，提出EP4CS框架——一种面向大语言模型代码摘要的增强型提示学习框架。首先设计Mapper模块，通过预训练&lt;代码，知识&gt;对促进提示向量根据语言模型输出进行优化更新；同时开发结构分析智能体（Struct-Agent），使语言模型能深入解析编程语言的语法结构以更精准理解复杂代码。实验结果表明：在相同参数规模下，本框架相较现有基线方法显著提升性能。基于StarCoderBase1B的Java测试中，EP4CS在BLEU、METEOR和ROUGE-L指标分别提升6.59%、7.06%与4.43%，同时展现出强劲的鲁棒性。在SentenceBERT语义评估维度更接近真实场景需求。人工评估与案例研究证实，EP4CS生成的摘要质量更高、相关性更强，全面优于基线方法。</span></span></p><p cid="n109" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728949" target="_blank">https://doi.org/10.1145/3728949</a></span></span></p><h3 cid="n110" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">32、Enhancing Smart Contract Security Analysis with Execution Property Graphs</span></span></h3><p cid="n111" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">智能合约漏洞已导致重大经济损失，随着其复杂性的日益增加，彻底防范黑客攻击变得愈发困难。这一趋势凸显了对高级取证分析和实时入侵检测的迫切需求，其中动态分析在剖析智能合约执行过程中发挥着关键作用。因此，亟需一种统一且通用的智能合约执行表示方法，并辅以高效的技术手段，以实现对各类新兴攻击的建模与识别。</span></span></p><p cid="n112" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们提出Clue——一个专为以太坊虚拟机设计的动态分析框架。其核心能力在于捕获合约执行过程中的关键运行时信息，并采用创新的基于图的表示方法：执行属性图。Clue的关键特性是其创新的图遍历技术，该技术擅长检测复杂攻击，包括（只读型）重入攻击和价格操纵攻击。评估结果表明，Clue以高真阳性率和低假阳性率展现出卓越性能，优于现有最先进工具。此外，Clue的高效性使其成为取证分析和实时入侵检测的双重利器。</span></span></p><p cid="n113" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728924" target="_blank">https://doi.org/10.1145/3728924</a></span></span></p><h3 cid="n114" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">33、Enhancing Vulnerability Detection via Inter-procedural Semantic Completion</span></span></h3><p cid="n115" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">受深度学习进展的启发，众多基于学习的漏洞检测方法应运而生，这些方法主要基于函数级操作以实现可扩展性。然而这种设计存在一个关键局限：许多漏洞跨越多个函数，导致函数级方法丢失被调用函数的语义信息，无法捕捉真实的漏洞模式。为解决这一问题，我们提出VulnSC框架，该创新框架通过补充过程间语义来增强基于学习的检测方法。VulnSC为数据集检索被调用函数的源代码，并利用大语言模型（LLMs）配合精心设计的提示词生成函数摘要。经过摘要增强的数据集被输入神经网络，以实现更精准的漏洞检测。VulnSC是首个将过程间语义整合到现有基于学习的漏洞检测方法中，同时保持可扩展性的通用框架。我们在两个广泛使用的数据集上对四种最先进的基于学习方法进行评估，实验结果表明VulnSC以极小的额外计算开销显著提升了检测性能。</span></span></p><p cid="n116" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728912" target="_blank">https://doi.org/10.1145/3728912</a></span></span></p><h3 cid="n117" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">34、Extended Reality Cybersickness Assessment via User Review Analysis</span></span></h3><p cid="n118" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">摘要近年来，扩展现实（XR）软件生态系统已成为下一代普适计算平台，其通过沉浸式交互体验为用户提供服务。然而XR生态系统存在晕眩症问题，会严重影响用户舒适度与安全，引发头痛、定向障碍等症状，这使得有效评估晕眩症成为亟待解决的重要课题。当前评估XR软件晕眩症的先进方法通常需在用户使用XR时监测其生理指标，这种方法严重依赖人工游戏测试，存在可扩展性受限的问题。XR应用商店中的用户评论能为开发者提供应用晕眩症评级及其成因的重要信息，但海量用户评论难以通过人工方式分析，且现有自动评论分析方法大多仅能提供粗粒度结果（如提取评论讨论的若干高层主题组）。大语言模型（LLM）的最新进展可能带来新机遇，但直接利用LLM评估XR晕眩症存在挑战：LLM对大量短文本处理效果不佳，且上下文窗口有限。为此，我们提出XRCare框架——通过细粒度用户评论分析实现XR应用晕眩症自动评估与根源推理的综合解决方案。该框架包含三阶段：（1）洞察池构建：汇集领域专家提供的晕眩症分析链及对应分析结果；（2）推理图谱构建：通过自演进层级图动态提取、分类和维护用户评论中引发晕眩症的成因；（3）多智能体演绎推理：利用多智能体系统模拟多样化用户群体，分析晕眩症强度等级并追溯成因。这种结构化方法使XRCare能系统化识别、分类和处理晕眩症实例。实验方面，我们构建了包含来自9,667款XR应用的685,111条用户评论的大规模数据集。评估表明，XRCare在最佳基线基础上将F1分数提升20.63%，平均超越所有基线32.27%，同时提供更精确细致的可解释性洞察。</span></span></p><p cid="n119" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728933" target="_blank">https://doi.org/10.1145/3728933</a></span></span></p><h3 cid="n120" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">35、FANDANGO: Evolving Language-Based Testing</span></span></h3><p cid="n121" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于语言的模糊测试器利用形式化输入规范（语言）为被测程序生成任意规模且多样化的有效输入集。现代基于语言的测试生成器结合语法与约束条件，以满足句法和语义层面的输入约束。该领域的领先输入生成工具ISLA采用符号约束求解技术处理输入约束。虽然使用求解器使ISLA成为精度最高的模糊测试器之一，但也导致其运行速度缓慢。</span></span></p><p cid="n122" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文探索以基于搜索的测试作为符号约束求解的替代方案。我们采用遗传算法，通过输入规范迭代生成候选输入，依据既定约束评估这些输入，通过句法有效的变异演化输入种群，保留适应度更优的个体直至满足语义输入约束。这种类似于自然遗传进化的演化过程，逐步产生能同时覆盖语义和句法的改进输入。此项改进显著提升了基于语言的测试效率：实验表明，相较于ISLA，我们基于搜索的FANDANGO原型在保持同等精度的前提下，速度提升一至三个数量级。</span></span></p><p cid="n123" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于搜索的方法不再将约束限制于约束求解器的（微型）语言范畴。FANDANGO允许约束条件使用完整的Python语言及其库。这种表达自由度为测试人员提供了前所未有的测试输入塑造灵活性，使其能够设定任意测试生成目标：&#34;请生成1000个有效测试输入，其中电压字段遵循高斯分布且始终不超过20毫伏&#34;。</span></span></p><p cid="n124" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://dl.acm.org/doi/10.1145/3728915" target="_blank">https://dl.acm.org/doi/10.1145/3728915</a></span></span></p><h3 cid="n125" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">36、Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models</span></span></h3><p cid="n126" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大型语言模型（LLM）虽在众多应用中展现出卓越性能，却会无意间吸收训练数据中的伪相关性，导致偏见概念与特定社会群体之间形成刻板印象关联。这些关联延续甚至放大了有害的社会偏见，引发了对公平性的重大关切——这正是软件工程领域的核心议题。为缓解此类偏见，现有研究尝试在推理过程中将模型嵌入投影至无偏见空间，但由于其与下游社会偏见的对齐程度较弱，这些方法效果有限。受到LLM中概念认知主要通过线性关联记忆机制（即MLP层中键值映射）实现的启发，我们提出偏见概念与社会群体同样以实体（键）和信息（值）对的形式编码，可通过操作这种编码机制促进更公平的关联。为此，我们提出公平中介器（FairMed）——一个高效且有效的偏见缓解框架，通过中和刻板印象关联来实现去偏。该框架包含两个核心组件：刻板关联探测器和对抗去偏中和器。探测器通过使用以偏见概念（键）为核心的提示词，捕获MLP层激活中编码的刻板关联，并检测社会群体（值）的发射概率；随后，对抗去偏中和器在推理过程中干预MLP激活，使不同社会群体间的关联概率趋于均衡。在九类受保护属性上的大规模实验表明，FairMed在效果上显著优于现有最优方法，平均偏见削减率最高达84.42%。</span></span></p><p cid="n127" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728881" target="_blank">https://doi.org/10.1145/3728881</a></span></span></p><h3 cid="n128" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">37、Finding 709 Defects in 258 Projects: An Experience Report on Applying CodeQL to Open-Source Embedded Software (Experience Paper)</span></span></h3><p cid="n129" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">嵌入式软件部署于全球数十亿设备中，包括医疗设备和自动驾驶汽车等安全关键系统。其缺陷可能造成严重后果。由于多数嵌入式软件产品整合了开源嵌入式软件（EMBOSS），采用适当机制规避缺陷对EMBOSS工程师至关重要。静态应用安全测试（SAST）工具作为常见安全实践手段，可帮助识别高频漏洞。现有SAST研究主要针对常规（非嵌入式）软件，缺乏对嵌入式软件领域的应用认知。嵌入式软件在语义结构、代码实践和构建配置方面与常规软件存在显著差异，这些因素都会影响SAST工具的实际效能。</span></span></p><p cid="n130" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文通过大规模实证研究，分析了258个主流EMBOSS项目的SAST应用现状。结合程序分析和开发者调研（N=25）发现：仅3%的项目采用超越基础编译器分析的高级SAST工具。开发者认为工具有效性不足和误报率高是主要制约因素。为此，我们应用最先进的CodeQL SAST工具进行实测，评估其易用性与实效性。在258个项目中，CodeQL检出709个真实缺陷（误报率34%），其中535个（75%）为潜在安全漏洞（涉及微软、亚马逊和阿帕奇基金会维护的重点项目）。EMBOSS工程师已确认376个缺陷（53%），主要通过合并我们提交的拉取请求；另促成2个CVE漏洞编号的分配。基于此，我们提议将检测流程集成至EMOSS持续集成（CI）管道，已有37个活跃仓库（占比71%）采纳该方案。研究表明：当代SAST工具具备低误报率与高缺陷检出效能，我们强烈建议EMBOSS工程师予以采用。</span></span></p><p cid="n131" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728923" target="_blank">https://doi.org/10.1145/3728923</a></span></span></p><h3 cid="n132" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">38、Fixing Outside the Box: Uncovering Tactics for Open-Source Security Issue Management</span></span></h3><p cid="n133" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">在快速演变的软件开发领域中，开源软件（OSS）安全漏洞的修复已变得至关重要。然而，学术界和工业界现有研究与工具主要依赖有限解决方案（如脆弱版本调整和采用补丁）来处理已识别的漏洞。而开源社区实际采用了更为灵活多样的应对策略，亟需通过整体性实证研究来探索这些多样化策略的普及程度、分布特征、偏好选择及实施效果。为此，本文对开源项目中的漏洞修复策略（RT）进行了系统分类研究，并评估了各类策略的优劣。本研究通过对GitHub上21,187个问题进行实证分析，揭示了开源社区修复策略的覆盖范围及有效性。我们构建了包含44种具体修复策略的层次化分类体系，并评估了其修复效果和实施成本。研究发现：社区高度依赖替代依赖库、漏洞规避等社区驱动策略（其中44%尚未被前沿工具支持），通过分析修复方案的采纳情况和拒绝原因，揭示了社区对特定修复方法的偏好。研究同时指出现代漏洞数据库存在严重缺陷——54%的CVE条目缺乏修复建议，而GitHub议题中93%的可操作解决方案可有效弥补这一缺口。</span></span></p><p cid="n134" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728977" target="_blank">https://doi.org/10.1145/3728977</a></span></span></p><h3 cid="n135" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">39、FreeWavm: Enhanced WebAssembly Runtime Fuzzing Guided by Parse Tree Mutation and Snapshot</span></span></h3><p cid="n136" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">WebAssembly作为一种低级、可移植的语言，已被广泛应用于浏览器和区块链等多个领域，成为推动互联网发展的革命性力量。然而，WebAssembly运行时中的缺陷和漏洞会在运行WebAssembly应用程序时导致意外结果。目前已有多种方案被提出用于检测WebAssembly运行时的漏洞，其中模糊测试因其显著效果成为最具前景和说服力的方法。尽管潜力巨大，但由于WebAssembly运行时语法复杂性——现有方法缺乏对独特模块化代码结构的深入理解，导致生成的测试输入难以触及运行时深层逻辑，限制了其揭示漏洞的有效性。</span></span></p><p cid="n137" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为弥补这一不足，我们提出FreeWavm——通过激进突变WebAssembly代码结构来模糊测试运行时的新型框架。技术层面，我们将WebAssembly字节码转换为能捕捉代码结构复杂特征的解析树格式。为生成具有意义的测试输入，设计了结构感知突变模块：采用定制化节点优先级策略筛选解析树中的关键节点，并施加特定结构突变。为确保突变后测试输入的有效性，FreeWavm配备自动修复机制来修补突变后的解析树。此外，我们利用解析树快照促进输入进化与整体模糊测试流程。通过在多类WebAssembly运行时上进行广泛实验，实证结果表明FreeWavm能有效触发运行时中结构特异性崩溃，性能优于同类方案。该框架已发现69个未知漏洞，其中24个目前已获得CVE编号。</span></span></p><p cid="n138" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728877" target="_blank">https://doi.org/10.1145/3728877</a></span></span></p><h3 cid="n139" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">40、Freesia: Verifying Correctness of TEE Communication with Concurrent Separation Logic</span></span></h3><p cid="n140" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">可信执行环境（TEE）作为现代处理器中的安全扩展，为敏感代码和数据提供了安全的运行时环境。尽管TEE旨在保护应用程序及其私有数据，但其庞大的代码库常存在可能危及数据安全的漏洞。虽然已有形式化验证工作针对TEE标准及实现的功能与安全性展开，但对并发场景下TEE正确性的验证仍不充分。本文提出一种名为Freesia的增强方案，通过形式化验证的并发分离逻辑确保TEE的并发安全性。基于对GlobalPlatform TEE标准的深入分析，Freesia解决了TEE通信接口中的数据竞争问题，并确保客户端与TEE间共享内存的一致性保护。我们在开源TEE平台OP-TEE中实现了Freesia原型，并利用Iris并发分离逻辑框架对其并发正确性进行建模与验证。通过实际案例研究和性能评估，进一步证明了Freesia的有效性与高效性。</span></span></p><p cid="n141" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728967" target="_blank">https://doi.org/10.1145/3728967</a></span></span></p><h3 cid="n142" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">41、GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific Bugs</span></span></h3><p cid="n143" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">图形用户界面（GUI）测试对于确保移动应用可靠性至关重要。现有最先进的GUI测试方法虽能成功探索更多应用场景并发现应用崩溃等通用缺陷，但工业级GUI测试还需检测特定于应用的缺陷，例如屏幕布局、控件位置或GUI转场效果与设计稿之间的偏差。这些由应用设计师创建的设计稿明确了预期屏幕、控件及其对应行为。验证GUI设计与实现的一致性虽耗时费力，却在工业GUI测试中具有重要作用。  </span></span></p><p cid="n144" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本研究提出一种检测移动应用设计与实现间不一致性的方法。移动设计通常包含两类设计稿：(1) 预期屏幕外观（如控件布局、颜色和形状）；(2) 预期屏幕行为（如带有文本描述的控件如何触发屏幕跳转）。给定设计稿及其对应应用实现，本方法可同时检测屏幕级和流程级不一致性。在屏幕检测方面，通过将屏幕抽象为控件容器（每个控件以位置、宽高和类型表示），定义控件偏序关系及替换、插入、删除操作的代价，将屏幕匹配问题转化为可优化的控件对齐问题。在流程检测方面，将指定GUI转场转化为屏幕操作序列（如点击、长按、文本输入），并提出视觉提示机制使视觉语言模型推断屏幕控件的具体操作，从而验证预期转场是否被正确实现。  </span></span></p><p cid="n145" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">在80个移动应用和160个设计稿上的实验表明：(1) 屏幕不一致性检测精度达99.8%，召回率达98.6%，分别较GVT等现有最优方法提升66.2%和56.6%；(2) 流程不一致性检测误差为零。此外，在交易类移动应用上的工业案例研究中，本方法成功检测出9个应用缺陷且均获原应用专家确认。代码已开源：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://github.com/code-philia/GUIPilot" target="_blank">https://github.com/code-philia/GUIPilot</a></span></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">。</span></span></p><p cid="n146" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728909" target="_blank">https://doi.org/10.1145/3728909</a></span></span></p><h3 cid="n147" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">42、Gamifying Testing in IntelliJ: A Replicability Study</span></span></h3><p cid="n148" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">游戏化是一种新兴技术，旨在提升传统枯燥任务（如软件测试）中的参与度和绩效。已有研究表明，通过提供成就和反馈机制，游戏化系统具有改善软件测试流程的潜力。然而，仍需在不同环境、编程语言和参与者群体中进一步验证这些益处。本文旨在复现并验证IntelliGame（一款IntelliJ IDEA游戏化插件）的效果，该插件旨在激励开发者编写和执行测试。研究目标是将早期研究中观察到的效益推广至新语境（即TypeScript编程语言和更大规模的参与者群体）。本次复现研究包含一项受控实验，招募174名参与者并分为两组：一组使用IntelliGame插件，另一组不使用任何游戏化插件。研究采用双组实验设计，比较两组在测试行为、覆盖率、变异分数及参与者反馈方面的差异。通过测试指标和参与者问卷收集数据，并进行统计分析以确定统计显著性。使用IntelliGame的参与者在测试实践中表现出比对照组更高的参与度和生产力，具体体现在创建更多测试用例、提高测试执行频率以及增强测试工具使用率。这些改进最终催生了更优质的代码实现，凸显了游戏化在提升功能成果和激励用户参与测试方面的有效性。本复现研究证实，通过IntelliGame实现的游戏化能对软件测试行为和开发者参与编码任务产生积极影响。这些发现表明，将游戏元素集成到测试环境中可成为改进软件测试实践的有效策略。</span></span></p><p cid="n149" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728983" target="_blank">https://doi.org/10.1145/3728983</a></span></span></p><h3 cid="n150" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">43、GoPV: Detecting Blocking Concurrency Bugs Related to Shared-Memory Synchronization in Go</span></span></h3><p cid="n151" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">Go是一种流行的并发编程语言，它采用消息传递和共享内存同步原语来实现不同线程（即goroutine）间的交互。然而，同步原语的误用极易导致阻塞型并发缺陷，包括死锁和goroutine泄漏。尽管与消息传递相关的阻塞型并发缺陷日益受到关注，但针对共享内存同步原语误用引发的阻塞型并发缺陷的研究却十分有限。本文提出GoPV——一个基于静态分析的并发缺陷检测工具，通过执行并发分析和（后）支配者分析来判定同步原语是否被误用，从而识别阻塞型并发缺陷。我们在8个基准测试程序和21个大型真实Go项目上对GoPV进行评估。实验结果表明，GoPV不仅成功检测出8个基准测试程序中所有与共享内存同步相关的阻塞型并发缺陷，还在2.78小时内从21个大型Go应用中发现了17个此类缺陷。</span></span></p><p cid="n152" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728979" target="_blank">https://doi.org/10.1145/3728979</a></span></span></p><h3 cid="n153" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">44、Hulk: Exploring Data-Sensitive Performance Anomalies in DBMSs via Data-Driven Analysis</span></span></h3><p cid="n154" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">性能对数据库管理系统（DBMS）至关重要，这类系统始终被设计用于高效处理不断变化的工作负载。然而，基于成本的优化器（CBO）及其交互机制的复杂性可能引发实现错误，导致数据敏感的性能异常。这些异常在某些数据集下可能导致与预期设计相比显著的性能下降。为诊断性能问题，DBMS开发者通常依赖直觉或与基线DBMS的执行时间进行对比，但这些方法忽略了数据集对性能的影响，导致仅能识别和解决部分性能问题。  </span></span></p><p cid="n155" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出Hulk系统，通过数据驱动分析自动探索这类数据敏感的性能异常。其核心思想是在数据集演化过程中识别性能异常：首先通过估算不同数据量下的合理响应时间范围来定位性能陡降点，随后通过寻找符合性能预期的合理执行计划来检测这些陡降点是否偏离预期性能。我们在六种广泛使用的DBMS（MySQL、MariaDB、Percona、TiDB、PostgreSQL和AntDB）上评估Hulk，共报告135个异常，其中129个被确认为新缺陷（含14个CVE漏洞），且94个属于数据敏感的性能异常。</span></span></p><p cid="n156" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728973" target="_blank">https://doi.org/10.1145/3728973</a></span></span></p><h3 cid="n157" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">45、ICEPRE: ICS Protocol Reverse Engineering via Data-Driven Concolic Execution</span></span></h3><p cid="n158" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">随着数字化转型的推进，工业控制系统（ICS）正变得日益开放和智能化。然而，ICS协议固有的脆弱性对设备和系统构成重大安全威胁。ICS协议的专有特性使得其安全分析和防护机制部署变得复杂。协议逆向工程旨在缺乏官方规范的情况下推断协议的语法、语义及状态机。传统协议逆向工程工具因缺乏可执行环境、推断策略不完善及网络流量质量低下而面临显著局限。本文提出ICEPRE——一种基于混合执行的新型数据驱动协议逆向工程方法，其独特地将网络轨迹与静态分析相结合。与传统依赖可执行环境的方法不同，ICEPRE通过静态追踪程序对特定输入消息的解析过程，采用创新的字段边界推断策略，通过分析协议解析器处理不同字段的方式推断协议语法。评估表明，ICEPRE在字段边界推断上显著优于现有工具：其F1分数达0.76、完美度分数0.67，而DynPRE、BinaryInferno、Nemeys和Netzob分别仅为（0.65, 0.35）、（0.42, 0.14）、（0.39, 0.09）和（0.27, 0.10）。这些结果印证了本方法卓越的整体性能。此外，ICEPRE在真实场景的专有协议测试中展现出优异表现，凸显了其在下游应用中的实用价值。</span></span></p><p cid="n159" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728982" target="_blank">https://doi.org/10.1145/3728982</a></span></span></p><h3 cid="n160" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">46、Identifying Multi-parameter Constraint Errors in Python Data Science Library API Documentation</span></span></h3><p cid="n161" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">现代人工智能与数据密集型软件系统高度依赖于数据科学和机器学习库，这些库提供核心算法实现和计算框架。这些库通过复杂API对外提供服务，其正确使用需遵循多个相互依赖参数间的约束条件。开发者通常需通过文档学习这些约束，任何偏差都可能导致意外行为。然而在API文档中保持多参数约束的正确性与一致性，仍是影响API兼容性和可靠性的重大挑战。为解决该问题，我们提出MPChecker工具，专门用于检测代码与文档间在多参数约束上的不一致性。该工具通过符号执行探索代码执行路径以识别代码级约束，并利用大语言模型（LLM）从文档中提取对应约束。我们提出定制化的模糊约束逻辑，以调和LLM输出的不确定性，并检测代码约束与文档约束间的逻辑不一致性。基于四个主流数据科学库构建的双数据集测试表明，MPChecker在126个不一致约束中成功识别117个，召回率达92.8%，有效验证了其检测能力。我们向库开发者提交了14个检测到的不一致问题，截至撰稿时已有11个获得确认。</span></span></p><p cid="n162" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728945" target="_blank">https://doi.org/10.1145/3728945</a></span></span></p><h3 cid="n163" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">47、Improving Deep Learning Framework Testing with Model-Level Metamorphic Testing</span></span></h3><p cid="n164" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">深度学习（DL）框架是DL软件系统的核心组件，其缺陷可能引发严重事故，因此需要有效的测试方法。现有研究通常采用DL模型或单一接口作为测试输入，通过分析执行结果来检测缺陷。然而浮点误差、固有随机性及测试输入的复杂性，使得执行结果分析面临巨大挑战，导致现有方法缺乏合适的测试预言。部分研究者采用蜕变测试应对该挑战，基于单一框架接口的输入数据和参数设置设计蜕变关系（MR），通过生成输出一致的等价测试输入来验证结果。这类方法虽具成效，仍存在三大局限：（1）现有MR忽视结构复杂性，限制测试输入多样性；（2）仅关注有限接口，制约泛化能力且需额外适配；（3）所检测缺陷多涉及单一接口结果一致性，难以发现多接口组合与运行时指标（如资源使用）相关的缺陷。为此，我们提出ModelMeta——一种面向DL框架的模型级蜕变测试方法，基于DL模型结构特性设计四种MR。该方法通过QR-DQN策略引导，利用多样化接口组合增强种子模型，生成输出一致的测试输入，并通过对训练损失/梯度、内存/GPU使用率及执行时间的细粒度分析来检测缺陷。我们在三大主流DL框架（MindSpore、PyTorch和ONNX）上使用涵盖图像分类至目标检测等十类实际任务的17个DL模型进行评估。结果表明：ModelMeta在测试覆盖率和生成测试输入多样性方面优于现有基线方法；共检测到31个新缺陷（其中27个获官方确认，11个已修复），包括7个现有方法无法检测的缺陷（5个资源使用错误和2个低效缺陷），证明了该方法的实用性。</span></span></p><p cid="n165" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728972" target="_blank">https://doi.org/10.1145/3728972</a></span></span></p><h3 cid="n166" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">48、Incremental Verification of Concurrent Programs through Refinement Constraint Adaptation</span></span></h3><p cid="n167" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">程序在其整个生命周期中持续演化。对每个版本从头开始进行验证通常不切实际，尤其是对于并发程序而言。设计高效的并发程序增量验证技术具有迫切需求。我们专注于面向并发程序验证的抽象精化技术。当程序被修改时，先前版本验证过程中生成的精化约束会被适配到新程序，以避免冗余分析。针对基于调度约束的抽象精化方法（当前最高效的并发程序验证精化方法之一），我们提出了基于内核源码的精化约束适配方案。本方法支持所有类型的程序修改，并能根据变更生成适配后的精化约束。在SV-COMP 2024基准测试集和Nidhugg基准测试上的评估表明，我们的方法取得了显著成效：实验中大多数先前版本验证生成的精化约束可成功适配至修改后的程序。与从头验证修改后程序相比，我们的增量验证方法对复杂程序可实现两个数量级的加速比。</span></span></p><p cid="n168" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728976" target="_blank">https://doi.org/10.1145/3728976</a></span></span></p><h3 cid="n169" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">49、Intention-Based GUI Test Migration for Mobile Apps using Large Language Models</span></span></h3><p cid="n170" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">图形用户界面（GUI）测试是移动应用质量保障的主要方法之一。人工构建高质量的GUI测试用例成本高昂且劳动密集，这推动了多种自动化方法的发展，这些方法旨在将测试用例从源应用迁移至目标应用。现有方法主要将该测试迁移任务视为控件匹配问题，在应用间交互逻辑保持一致时表现良好。但当不同应用对特定功能存在交互逻辑差异时（这是跨应用的常见场景），现有方法则面临挑战。为解决这一局限，本文提出了一种名为ITeM的新型测试迁移方法。与将问题建模为控件匹配任务的现有工作不同，ITeM通过采用具备大语言模型理解与推理能力的双阶段框架开辟了新路径：首先通过过渡感知机制生成测试意图，其次通过基于动态推理的机制实现这些意图。该方法能有效应对源应用与目标应用间交互逻辑的差异。在35个真实安卓应用上开展的280项测试迁移任务实验表明，ITeM相比最先进方法具有显著优越的有效性和效率。</span></span></p><p cid="n171" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728978" target="_blank">https://doi.org/10.1145/3728978</a></span></span></p><h3 cid="n172" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">50、KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity Analysis</span></span></h3><p cid="n173" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">二进制代码相似性分析（BCSA）是网络安全等众多领域的关键研究方向。其中，函数级差异比对工具在BCSA中应用最为广泛：它们通过逐函数匹配来评估二进制程序间的相似性。然而此类方法具有较高的时间复杂度，难以适应大规模场景（如1对n或n对n搜索）。为实现高效且精准的程序级BCSA，我们提出KEENHash——一种通过大语言模型（LLM）生成函数嵌入的新型哈希方法，将二进制代码转换为程序级表征。该方法结合K-Means聚类和特征哈希技术，将二进制程序压缩为紧凑的定长程序嵌入，从而实现了高效的大规模程序级BCSA，其性能超越现有最优方法。实验结果表明：在保持精度的前提下，KEENHash比最先进的函数匹配工具快至少215倍。在53亿次相似性比对的大规模场景中，KEENHash仅需395.83秒，而传统工具至少需要56天。我们在包含202,305个二进制程序的大规模数据集上进行程序克隆搜索测试，与4种前沿方法相比，KEENHash以至少23.16%的优势全面胜出，并在恶意软件检测的大规模BCSA安全场景中展现出显著优越性。</span></span></p><p cid="n174" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728911" target="_blank">https://doi.org/10.1145/3728911</a></span></span></p><h3 cid="n175" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">51、KRAKEN: Program-Adaptive Parallel Fuzzing</span></span></h3><p cid="n176" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">并行模糊测试通过利用多核计算机加速测试过程，已在工业级软件缺陷检测领域获得广泛应用。然而由于静态推断模糊测试运行时存在困难，针对不同特性的程序制定高效并行策略仍具挑战性。现有方案仍采用预定义策略应对不同程序，导致性能未达最优。</span></span></p><p cid="n177" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出KraKen——一种新型程序自适应并行模糊测试器，通过动态策略优化提升测试效率。其核心在于：通过代码覆盖率变化等运行时反馈可观测并行模糊测试的低效现象，从而调整策略以避免低效路径搜索，逐步逼近最优方案。基于此，我们将寻找最优策略的任务构建为优化问题，通过动态最大化特定目标函数，逐步适配出针对具体程序的最佳策略。</span></span></p><p cid="n178" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们在C/C++中实现了KraKen，并在19个真实世界程序上与6种最先进并行模糊测试器进行对比评估。实验结果表明，KraKen在给定时间内可实现54.7%的代码覆盖率提升，并多发现70.2%的程序错误。此外，KraKen已在37个热门开源项目中发现192个错误，其中119个被分配了CVE编号。</span></span></p><p cid="n179" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728882" target="_blank">https://doi.org/10.1145/3728882</a></span></span></p><h3 cid="n180" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">52、LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation</span></span></h3><p cid="n181" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">代码生成旨在根据输入需求自动生成代码，显著提升开发效率。基于大型语言模型（LLM）的最新方法已展现出突破性成果，彻底改变了代码生成任务。尽管性能前景可观，LLM生成的代码常存在幻觉现象，尤其在需要处理实际开发过程中复杂上下文依赖的代码生成场景中。虽然已有研究分析了LLM代码生成中的幻觉问题，但其研究范围局限于独立函数生成。本文通过实证研究，在更贴近实际且复杂度更高的仓库级代码生成场景中，系统探究LLM幻觉的现象、机理与缓解策略。首先，我们人工检测了六种主流LLM的代码生成结果，建立了LLM生成代码的幻觉分类体系；继而详细阐述了幻觉现象特征，并分析了不同模型间的幻觉分布规律；随后深入剖析幻觉成因，识别出四大潜在致幻因素；最后提出基于检索增强生成（RAG）的缓解方法，该方法在所有研究的LLM中均展现出持续有效的改善效果。</span></span></p><p cid="n182" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728894" target="_blank">https://doi.org/10.1145/3728894</a></span></span></p><h3 cid="n183" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">53、LLM4SZZ: Enhancing SZZ Algorithm with Context-Enhanced Assessment on Large Language Models</span></span></h3><p cid="n184" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">SZZ算法是识别缺陷引入提交（bug-inducing commits）的核心技术，为许多软件工程研究（如缺陷预测和静态代码分析）提供基础，从而提升软件质量并优化维护实践。自该算法提出以来，研究者已开发多种改进版本以增强其性能。大多数改进依赖静态技术或启发式假设，虽易于实现，但性能提升有限。近期出现了一种基于深度学习的SZZ算法，但其需复杂预处理且仅支持单一编程语言；虽提高了精确率，却降低了召回率。此外，现有改进大多忽略关键信息（如提交消息和补丁上下文），且仅适用于涉及代码删除行的缺陷修复提交。  </span></span></p><p cid="n185" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大语言模型（LLMs）的出现为解决这些问题提供了新机遇。本研究系统分析了LLMs的优势与局限，提出LLM4SZZ框架，采用两种方法（基于排序的识别和上下文增强识别）处理不同类型的缺陷修复提交。我们根据LLM对缺陷的理解能力及其判断提交是否包含缺陷的能力来选择方法：上下文增强识别为LLM提供更丰富的上下文，要求其从候选提交中定位缺陷引入提交；基于排序的识别则让LLM从缺陷修复提交中筛选缺陷代码语句，并按其与根本原因的相关性排序。实验结果表明，LLM4SZZ在三个数据集上均优于所有基线模型，F1分数提升6.9%至16.0%，且未显著牺牲召回率。此外，LLM4SZZ能识别基线模型未能检测的缺陷引入提交，占比分别达三个数据集提交总量的7.8%、7.4%和2.5%。</span></span></p><p cid="n186" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728885" target="_blank">https://doi.org/10.1145/3728885</a></span></span></p><h3 cid="n187" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">54、LogBase: A Large-Scale Benchmark for Semantic Log Parsing</span></span></h3><p cid="n188" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大规模软件系统生成的日志包含大量有用信息。作为自动化日志分析的第一步，日志解析技术已被广泛研究。通用日志解析方法主要关注从原始日志中识别静态模板，但忽略了动态日志参数中隐含的更重要的语义信息。随着智能运维（AIOps）的普及，传统日志解析方法已无法满足各类下游任务的需求。研究者开始探索新一代日志解析技术——语义日志解析，旨在同时识别日志模板和参数语义。然而，现有数据集中语义标注的缺失阻碍了语义日志解析器的训练与评估，制约了该领域的发展。  </span></span></p><p cid="n189" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为填补这一空白并推动语义日志解析研究，我们构建了首个语义日志解析基准数据集LogBase。该数据集涵盖130个热门开源项目的日志，包含85,300条带语义标注的日志模板，在日志源多样性和模板丰富性上均超越现有数据集。为实现LogBase的构建，我们开发了语义日志解析数据集构建框架GenLog。该框架从GitHub热门开源仓库中挖掘日志模板-参数-上下文三元组，并采用思维链（CoT）技术驱动大语言模型（LLMs）生成高质量日志。同时，GenLog通过人工反馈优化生成数据质量并确保其可靠性。该框架具备高度自动化与成本效益，可助力研究者高效构建语义日志解析数据集。此外，我们还为LogBase设计了一套综合评估指标，涵盖通用日志解析器指标、语义日志解析器专项指标及基于LLM的解析器指标。  </span></span></p><p cid="n190" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于LogBase，我们对15种现有日志解析器进行了全面评估，揭示了它们在复杂场景下的真实性能。我们相信，这项工作将为研究者提供宝贵数据、可靠工具和深入见解，以支持并引导语义日志解析的未来研究。</span></span></p><p cid="n191" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728969" target="_blank">https://doi.org/10.1145/3728969</a></span></span></p><h3 cid="n192" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">55、MLLM-Based UI2Code Automation Guided by UI Layout Information</span></span></h3><p cid="n193" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">将用户界面转换为代码（UI2Code）是网站开发中的关键环节，该过程耗时且费力。实现UI2Code的自动化对提升开发效率至关重要。现有基于深度学习的方法严重依赖大量标注训练数据，且难以泛化到真实世界中未见过的网页设计。多模态大语言模型（MLLMs）的出现为解决该问题提供了可能，但其难以理解UI中的复杂布局并生成保留布局的精确代码。为此，我们提出LayoutCoder——一种基于MLLM的创新框架，可从真实网页图像生成UI代码，包含三个核心模块：（1）元素关系构建：通过识别和分组具有相似结构的组件来捕捉UI布局；（2）UI布局解析：生成UI布局树以指导后续代码生成；（3）布局引导的代码融合：生成保留布局的精确代码。为进行评估，我们构建了包含350个真实网站的新基准数据集Snap2Code（分为可见与不可见部分以缓解数据泄露问题），并采用流行数据集Design2Code。大量实验表明，LayoutCoder在所有数据集上平均BLEU分数提升10.14%，CLIP分数提升3.95%，显著优于现有最优方法。</span></span></p><p cid="n194" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728925" target="_blank">https://doi.org/10.1145/3728925</a></span></span></p><h3 cid="n195" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">56、MoDitector: Module-Directed Testing for Autonomous Driving Systems</span></span></h3><p cid="n196" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">测试自动驾驶系统（ADS）对于确保其安全性、可靠性和性能至关重要。尽管现有多种测试方法能够生成多样化的高难度场景以发现潜在漏洞，但这些方法通常将ADS视为黑盒，主要关注识别系统级故障（如碰撞或险兆事故），而无法定位导致故障的具体模块。这种对故障根本原因的认知缺失阻碍了有效的调试与后续系统修复。此外，现有方法在生成能够从系统层面充分测试ADS各独立模块（如感知、预测、规划与控制）的违规场景方面存在不足。为弥补这一缺陷，我们提出MoDitector——一种具备根本原因识别能力的ADS测试方法，该方法可生成专门针对目标ADS模块弱点设计的安全关键场景。与现有方法不同，MoDitector不仅能产生导致违规的场景，还能精确定位引发每个故障的具体责任模块。具体而言，我们通过引入模块专用预言机（Module-Specific Oracles）自动检测模块级错误，并识别导致系统级违规的根本原因模块。为有效生成模块专属故障，我们提出一种模块导向测试策略，该策略融合模块专用反馈与自适应场景生成技术来指导测试过程。我们在四个关键ADS模块和四个代表性测试场景中评估MoDitector。实验结果表明，MoDitector能高效生成可归因于特定目标模块的故障场景，总计生成216.7个预期场景，显著优于最佳基线方法（仅生成79.0个场景）。本研究通过聚焦系统内模块专属错误的识别与修正，突破了传统黑盒故障检测的局限，代表了ADS测试领域的重大创新。</span></span></p><p cid="n197" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728876" target="_blank">https://doi.org/10.1145/3728876</a></span></span></p><h3 cid="n198" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">57、Model Checking Guided Incremental Testing for Distributed Systems</span></span></h3><p cid="n199" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">近年来，模型检测引导测试（MCGT）方法被提出用于系统化测试分布式系统。该方法通过遍历从分布式系统形式化规范导出的完整已验证抽象状态空间来自动生成测试用例，并检查目标系统在测试过程中是否表现正确。尽管MCGT具有有效性，但使用该技术测试分布式系统通常成本高昂且可能耗时数周。当分布式系统发生演进（如引入新功能或修复缺陷）时，这种低效问题会进一步加剧。我们必须为演进后的系统重新运行完整测试流程以验证其正确性，这使得MCGT不仅资源密集且效率低下。</span></span></p><p cid="n200" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为降低分布式系统演进过程中模型检测引导测试的开销，我们提出iMocket——一种新型的模型检测引导增量式测试方法。我们首先从形式化规范和系统实现中提取变更内容，随后识别抽象状态空间中受影响的状态，并专门针对这些状态生成增量测试用例，从而避免对未受影响状态的冗余测试。基于三个主流分布式系统的12个真实变更场景进行评估实验，结果表明：iMocket平均可减少74.83%的测试用例数量，并将测试时间降低22.54%至99.99%，显著证明了其在降低分布式系统测试成本方面的有效性。</span></span></p><p cid="n201" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728883" target="_blank">https://doi.org/10.1145/3728883</a></span></span></p><h3 cid="n202" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">58、More Effective JavaScript Breaking Change Detection via Dynamic Object Relation Graph</span></span></h3><p cid="n203" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">JavaScript库以其广泛使用、频繁的代码变更及对向后不兼容变更的高容忍度而著称。意识到这类破坏性变更可帮助开发者适应版本更新并规避负面影响。JavaScript社区已有多种专门或可用于检测破坏性变更的工具，但这些工具采用不同检测方式，且目前缺乏对这些方法的系统性综述。通过对流行JavaScript库的初步研究，我们发现现有方法（包括简单回归测试、基于模型的测试和类型差异分析）不仅会遗漏大量破坏性变更，还会产生大量误报。本文讨论了漏检与误报的产生原因。</span></span></p><p cid="n204" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于研究发现的启示，我们提出名为Diagnose的新方法：该方法通过API探索和基于强制执行的类型分析迭代构建对象关系图，随后对图谱进行精细化处理并在库的新版本中重构图谱以检测破坏性变更。通过在实证研究相同库集上的评估，Diagnose能检测出更多破坏性变更（60.2%）且误报更少，因此具备实际应用价值。</span></span></p><p cid="n205" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728980" target="_blank">https://doi.org/10.1145/3728980</a></span></span></p><h3 cid="n206" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">59、NADA: Neural Acceptance-Driven Approximate Specification Mining</span></span></h3><p cid="n207" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">仅从期望的软件行为（即正例）中挖掘高质量的有限状态自动机（FSA）十分困难，这源于搜索空间爆炸以及缺乏非期望软件行为（即反例）导致的过度泛化问题。为解决过度泛化问题，我们建议将该问题建模为从含噪声的正例与反例中搜索近似FSA，其中噪声源自用于拒绝过度泛化结果的合成反例。为在爆炸性搜索空间中获取有效的搜索偏置，我们将FSA接受度与神经网络推理相融合。核心贡献在于设计了一种神经网络，其参数分配对应于FSA，且其名为&#34;神经接受&#34;的推理过程能够模拟FSA接受行为。神经接受机制提供了一种高效量化FSA与噪声数据拟合程度的方法。我们提出NADA——一种神经接受驱动的搜索方法，通过接受正例与拒绝合成反例来指导近似FSA的搜索。NADA基于FSA离散搜索空间的连续松弛化改造及高效的梯度下降搜索算法实现。实验结果表明：相较于最先进方法，NADA显著提升了挖掘FSA的质量（平均提升41.63%的F1分数），且其搜索速度比挖掘次高质量FSA的方法快19.8倍。</span></span></p><p cid="n208" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728956" target="_blank">https://doi.org/10.1145/3728956</a></span></span></p><h3 cid="n209" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">60、No Bias Left Behind: Fairness Testing for Deep Recommender Systems Targeting General Disadvantaged Groups</span></span></h3><p cid="n210" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">推荐系统在现代社会中扮演着日益重要的角色，它们驱动着数字平台对新闻、音乐、职位列表等多种内容进行个性化推荐，深刻影响着日常生活的诸多方面。为提升个性化效果，这些系统常使用人口统计信息。然而，确保跨人口群体推荐质量的公平性仍具挑战性，尤其因为推荐系统易受用户反馈循环中&#34;富者愈富&#34;的马太效应影响。随着深度学习算法的普及，公平性问题的识别变得愈发复杂。研究者已开始探索利用优化算法识别最弱势用户群体的方法。尽管如此，次优弱势群体的研究仍显不足，这导致马太效应引发的偏见放大风险未被有效解决。本文主张同时识别最弱势与次优弱势群体的必要性，并提出基于自适应采样的FairAS方法实现该目标。通过对四个深度推荐系统和六个数据集的评估，FairAS在最弱势群体识别上相比当前最优公平性测试方法（FairRec）平均提升19.2%，同时将测试时间降低43.07%。此外，FairAS发现的额外次优弱势群体有助于提升系统公平性，在所有实验对象上相比FairRec平均实现70.27%的改进。</span></span></p><p cid="n211" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728948" target="_blank">https://doi.org/10.1145/3728948</a></span></span></p><h3 cid="n212" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">61、OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution</span></span></h3><p cid="n213" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">GitHub问题解决任务旨在自动处理代码仓库中报告的问题。随着大语言模型（LLM）的发展，该任务日益受到关注，多个基准测试被提出以评估LLM的问题解决能力。然而，现有基准存在三个主要局限：首先，当前基准仅关注单一编程语言，限制了跨语言仓库问题的评估；其次，它们通常覆盖领域范围狭窄，难以体现现实问题的多样性；第三，现有基准仅依赖问题描述中的文本信息，忽略了图像等多模态信息。本文提出OmniGIRL——一个多语言、多模态、多领域的GitHub问题解决基准，包含从四种编程语言（Python、JavaScript、TypeScript和Java）和八个不同领域的仓库中收集的959个任务实例。评估表明，当前LLM在OmniGIRL上表现有限，性能最佳的GPT-4o仅解决8.6%的问题。此外，我们发现现有LLM难以处理需要理解图像的问题，Claude-3.5-Sonnet在含图像信息的问题上仅解决10.5%。最后，我们分析了LLM在OmniGIRL上失败的原因，为未来改进提供见解。</span></span></p><p cid="n214" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728871" target="_blank">https://doi.org/10.1145/3728871</a></span></span></p><h3 cid="n215" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">62、OpDiffer: LLM-Assisted Opcode-Level Differential Testing of Ethereum Virtual Machine</span></span></h3><p cid="n216" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">随着以太坊的持续繁荣，以太坊虚拟机（EVM）已成为支撑数百万活跃智能合约的基石。直观而言，EVM中的安全问题可能导致智能合约间的行为不一致，甚至造成整个区块链网络的拒绝服务。然而据我们所知，目前仅有有限的研究聚焦于EVM安全性，且存在两大局限：1）测试输入多样性不足且缺乏无效语义；2）无法自动识别漏洞并定位根本原因。为弥补这一空白，我们提出OpDiffer——一种基于差分测试的EVM检测框架，通过结合大语言模型（LLM）与静态分析方法解决上述问题。我们开展了最大规模的评估实验，覆盖九种EVM实现，发现26个此前未知的漏洞（其中22个获开发者确认，3个获得CNVD编号）。相比最先进的基线方法，OpDiffer最高可分别提升71.06%、148.40%和655.56%的代码覆盖率。通过对实际部署的以太坊合约分析，我们预估7.21%的合约在特定环境配置下可能触发已发现的EVM漏洞，这将对以太坊生态系统产生严重的负面影响。</span></span></p><p cid="n217" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728946" target="_blank">https://doi.org/10.1145/3728946</a></span></span></p><h3 cid="n218" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">63、PatchScope: LLM-Enhanced Fine-Grained Stable Patch Classification for Linux Kernel</span></span></h3><p cid="n219" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">稳定的补丁分类在Linux内核漏洞管理中起着至关重要的作用，对长期支持（LTS）版本的稳定性和安全性具有重大意义。尽管现有工具能有效辅助判断补丁是否应合并至稳定版本，但无法确定哪些稳定补丁应被合并到哪些LTS版本中。该过程仍需发行版社区维护者根据各自版本需求进行人工筛选。  </span></span></p><p cid="n220" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为解决这一问题，我们提出PatchScope，其旨在预测补丁的具体合并状态。PatchScope包含两个组件：补丁分析与补丁分类。补丁分析利用大语言模型（LLMs），通过提交信息与代码变更生成详细的补丁描述，从而深化模型对补丁的语义理解；补丁分类采用预训练语言模型提取补丁的语义特征，并利用两阶段分类器预测补丁的合并状态。通过动态加权损失函数优化模型，以处理数据不平衡问题并提升整体性能。  </span></span></p><p cid="n221" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">鉴于当前主要维护Linux内核5.10与6.6版本，我们基于这两个版本进行了对比实验。实验结果表明，PatchScope能有效预测补丁的合并状态。</span></span></p><p cid="n222" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728944" target="_blank">https://doi.org/10.1145/3728944</a></span></span></p><h3 cid="n223" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">64、Pepper: Preference-Aware Active Trapping for Ransomware</span></span></h3><p cid="n224" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">勒索软件通过加密受感染系统中的文件并索要高额解密赎金，对企业与个人构成严重威胁。然而现有方法未能捕捉不同勒索软件家族的加密偏好，缺乏高效系统的主动防御方案。本文提出Pepper——一种基于偏好感知的主动式勒索软件诱捕方法，涵盖诱饵文件生成、部署与监控环节。通过对大量勒索软件家族的分析，我们识别出两种普遍存在的加密偏好：文件类型偏好与加密路径偏好。在勒索软件偏好的路径中部署符合其加密偏好的诱饵文件，能够为高效早期诱捕提供可能。Pepper融合基于图神经网络的推荐模型与专家知识，揭示不同勒索软件家族的文件与路径加密偏好，指导诱饵文件的生成与部署。此外，系统设计诱饵文件监控模块持续追踪文件变化并及时响应异常。大规模实验表明，Pepper实现98.68%的勒索软件检测率，平均仅损失2.27个文件，且在检测未知勒索软件变种时表现出强鲁棒性，同时不会干扰正常用户操作。</span></span></p><p cid="n225" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728932" target="_blank">https://doi.org/10.1145/3728932</a></span></span></p><h3 cid="n226" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">65、Porting Software Libraries to OpenHarmony: Transitioning from TypeScript or JavaScript to ArkTS</span></span></h3><p cid="n227" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">OpenHarmony正崛起为移动应用领域的重要力量，有望与行业巨头比肩。其主力开发语言ArkTS基于TypeScript(TS)和JavaScript(JS)进行强化，通过严格类型系统提升性能。生态建设需要开发者将主流TS/JS库移植至OpenHarmony，官方虽提供详细移植指南，但要求开发者深度掌握ArkTS语法规范、遵循移植规则并实施人工代码改造，因此自动化移植工具对提升效率和完善软件生态至关重要。  </span></span></p><p cid="n228" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">作为新兴编程语言，ArkTS目前缺乏支持自动化库移植的分析工具。而大语言模型(LLMs)的兴起为自动化移植任务提供了新机遇。基于LLM实现TS/JS库到OpenHarmony的自动化移植面临两大挑战：(1)LLMs对ArkTS代码接触有限，难以掌握其与JS/TS的语法差异及多样化适配场景；(2)项目级代码适配需修正大量语法不匹配问题，不同不匹配项与互依代码间的交互作用更增加了LLM的处理复杂度。  </span></span></p><p cid="n229" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为此，我们提出项目级自动代码适配方案ArkAdapter：针对挑战一，通过构建包含多场景专家经验的真实代码适配案例库，建立ArkTS语法理解知识库，借助小样本学习增强LLMs的适配能力；针对挑战二，基于依赖结构和语法不匹配代码粒度制定适配优先级策略，避免不同语法不匹配项及其关联代码的相互干扰。实验表明ArkAdapter在JS/TS库到ArkTS的移植中达到86.84%的高准确率。</span></span></p><p cid="n230" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728941" target="_blank">https://doi.org/10.1145/3728941</a></span></span></p><h3 cid="n231" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">66、Preventing Disruption of System Backup against Ransomware Attacks</span></span></h3><p cid="n232" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">摘要近年来，勒索软件对软件生态系统的威胁迅速增长。尽管已有深入研究，但新型勒索软件变种不断涌现，旨在规避现有基于加密的检测机制。本文提出Remembrall——通过监控并防止系统备份中断来防御勒索软件的新方案。该工具聚焦Windows系统卷影副本（VSC）的删除操作，捕获相关恶意事件并实时识别所有勒索软件痕迹。为确保全面防护，我们系统性地分类研究了应用层、操作系统层和硬件层中所有可能用于删除VSC的攻击行为。基于此分析，Remembrall通过检索系统事件信息实现精准识别，确保零漏报率。通过对最新勒索软件样本的评估，Remembrall在60个勒索软件家族检测中的F1分数比现有顶级反勒索软件工具提高4.31%-87.55%，并在实验中成功检测出8个零日勒索软件样本。</span></span></p><p cid="n233" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728880" target="_blank">https://doi.org/10.1145/3728880</a></span></span></p><h3 cid="n234" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">67、Productively Deploying Emerging Models on Emerging Platforms: A Top-Down Approach for Testing and Debugging</span></span></h3><p cid="n235" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">尽管现有的机器学习（ML）框架主要针对成熟平台（如在服务器级GPU上运行CUDA），但越来越多的需求希望在各种新兴场景中实现人工智能应用，例如在浏览器和移动端运行大语言模型（LLM）。然而，由于模型快速迭代以及新平台（如Metal和WebGPU）缺乏成熟的工具链和实践经验，在这些平台上部署新兴模型面临显著的软件工程挑战。传统的ML模型部署通常采用自下而上的方式：工程师先实现单个必需算子，再进行组合集成。但这种开发模式难以满足新兴ML应用的部署效率要求，其中测试与调试环节成为瓶颈。为此，我们提出TapML——一种自上而下的方法，旨在简化跨平台模型部署流程。传统自下而上方法需手动编写测试用例，而TapML通过算子级测试切分自动生成高质量的真实测试数据；此外，采用基于迁移的策略逐步将模型实现从成熟源平台转移至目标平台，最大限度缩小复合错误的调试范围。TapML已作为MLC-LLM项目的默认开发方法用于部署新兴ML模型。两年内，依托该方法成功在5个新兴平台上部署了涵盖27种模型架构的105个新兴模型。实践表明TapML在保证部署质量的同时显著提升了开发效率，并基于实际开发经验总结了全面案例研究，为新兴ML系统开发提供了最佳实践指南。</span></span></p><p cid="n236" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728957" target="_blank">https://doi.org/10.1145/3728957</a></span></span></p><h3 cid="n237" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">68、Program Analysis Combining Generalized Bit-Level and Word-Level Abstractions</span></span></h3><p cid="n238" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">抽象解释被广泛用于确定程序的数值属性。然而，当前抽象域主要关注数学语义，未能完全捕捉依赖机器整数语义且涉及大量位向量操作的实际程序的复杂性。本文提出了一种结合位级抽象和字级抽象的解决方案来捕捉机器整数语义。首先，我们通过补充所有必需操作作为标准抽象域，推广了Linux eBPF验证器中用于确定实际程序已知位与未知位的位级抽象。基于此抽象，我们设计了一个具备符号感知能力、同时保留上述位级和字级边界信息的抽象域。这两个层级的信息通过标准缩减积操作进行协作，以提高分析精度。我们在Crab分析器和内核外eBPF验证器PREVAL中实现了所提出的抽象域。实验证明其在分析SV-COMP基准程序、辅助硬件设计以及eBPF验证方面的有效性。</span></span></p><p cid="n239" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728905" target="_blank">https://doi.org/10.1145/3728905</a></span></span></p><h3 cid="n240" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">69、Program Feature-Based Benchmarking for Fuzz Testing</span></span></h3><p cid="n241" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">模糊测试是一种强大的软件测试技术，以其在识别软件漏洞方面的有效性而闻名。传统的模糊测试评估通常关注模糊测试工具在一组目标程序上的整体性能，但很少有基准测试考虑细粒度程序特征如何影响模糊测试效果。为弥补这一空白，我们提出了FeatureBench——一种新颖的基准测试框架，能通过可配置的细粒度程序特征生成测试程序，以增强模糊测试评估效果。通过系统回顾25项近期灰盒模糊测试研究，我们提取出7个可能影响测试性能的控制流与数据流相关程序特征。基于这些特征，我们生成了包含153个程序的基准测试集，并通过10个细粒度可配置参数进行控制。使用该基准测试集对11种模糊测试工具进行评估（每种工具均代表特定改进方向或是广泛使用的基准方案），结果表明：测试工具性能会随程序特征及其强度呈现显著差异，这凸显了将程序特性纳入模糊测试评估体系的重要性。</span></span></p><p cid="n242" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728899" target="_blank">https://doi.org/10.1145/3728899</a></span></span></p><h3 cid="n243" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">70、QTRAN: Extending Metamorphic-Oracle Based Logical Bug Detection Techniques for Multiple-DBMS Dialect Support</span></span></h3><p cid="n244" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">蜕变测试是一种广泛用于检测数据库管理系统（DBMS）中逻辑缺陷的方法，本文称为MOLT（基于蜕变关系的逻辑缺陷检测技术）。该技术通过构建SQL语句对（包括原始查询和变异查询），并评估执行结果是否符合预定义的蜕变关系来识别逻辑缺陷。然而，现有的MOLT严重依赖特定DBMS的语法生成有效SQL语句对，导致难以适配具有不同语法结构的各类DBMS。因此，当前仅支持少数主流DBMS（如PostgreSQL、MySQL和MariaDB），扩展至其他系统需大量人工投入。鉴于许多DBMS仍缺乏充分测试，亟需一种能够轻松扩展MOLT至异构DBMS的方法。  </span></span></p><p cid="n245" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出QTRAN——一种基于大语言模型（LLM）的新方法，可自动将现有MOLT扩展至多种DBMS。核心思路是利用LLM将现有MOLT中的SQL语句对翻译为目标DBMS的语法以进行蜕变测试。针对LLM对方言差异和蜕变机制理解有限的挑战，我们提出包含转换阶段和变异阶段的两阶段方法。QTRAN借鉴开发者创建MOLT的过程：通过理解目标DBMS语法生成原始查询，并利用定制化变异器执行突变。转换阶段通过识别潜在方言并利用SQL文档信息增强查询检索，使LLM能精准跨DBMS翻译原始查询；变异阶段则收集现有MOLT的SQL语句对微调预训练模型，专门适配变异任务，再使用定制化LLM对翻译后的原始查询进行变异，保留蜕变测试所需的定义关系。  </span></span></p><p cid="n246" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们将该方法实现为工具，并应用于扩展四种前沿MOLT至八种DBMS：MySQL、MariaDB、TiDB、PostgreSQL、SQLite、MonetDB、DuckDB和ClickHouse。评估结果表明，QTRAN转换的SQL语句对超过99%满足测试所需的蜕变关系，且在这些DBMS中检测到24个逻辑缺陷，其中16个被确认为独特的新缺陷。我们相信QTRAN的通用性将显著提升DBMS的可靠性。</span></span></p><p cid="n247" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728908" target="_blank">https://doi.org/10.1145/3728908</a></span></span></p><h3 cid="n248" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">71、Quantum Concolic Testing</span></span></h3><p cid="n249" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文首次提出了一个专为量子程序设计的混成测试框架。该框架针对量化量子态的量子控制语句提出了量子约束生成方法，并为量子变量提供了符号化方法。基于此框架，我们为量子程序的每条具体执行路径生成路径约束。这些约束条件指导新路径的探索，通过量子约束求解器确定结果以生成新颖的输入样本，从而提升分支覆盖率。本框架已在Python中实现并与Qiskit集成进行实践评估。实验结果表明，我们的混成测试框架不仅能提高分支覆盖率，还能生成高质量量子输入样本并检测程序缺陷，证明了其在量子编程和错误检测方面的有效性与高效性。在分支覆盖率方面，本框架对5量子位以下的量子程序实现了超过74.27%的覆盖效果。</span></span></p><p cid="n250" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728926" target="_blank">https://doi.org/10.1145/3728926</a></span></span></p><h3 cid="n251" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">72、REACCEPT: Automated Co-evolution of Production and Test Code Based on Dynamic Validation and Large Language Models</span></span></h3><p cid="n252" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">生产代码与测试代码的同步（称为PT协同演化）对软件质量至关重要。鉴于其涉及大量手动工作，研究人员尝试使用预定义启发式规则和机器学习模型实现PT协同演化的自动化。然而现有解决方案仍不完善：大多数方法仅能检测并标记过时测试用例，仍需开发者手动更新；同时现有方案准确率较低，尤其在真实软件项目中表现不佳。本文提出ReAccept——一种融合大型语言模型（LLM）、检索增强生成（RAG）和动态验证的新方法，以高精度实现全自动PT协同演化。ReAccept采用经验引导方法生成提示模板，用于识别和更新过程；在更新测试用例后，通过语法检查、语义验证和测试覆盖评估进行动态验证；若验证失败，则利用错误消息迭代优化补丁。为评估ReAccept的有效性，我们在包含537个Java项目的数据集上开展广泛实验，并与多种先进方法对比。结果表明，ReAccept在正确识别的过时测试代码上达到60.16%的更新准确率，较最优技术CEPROT提升90%。这些发现证明ReAccept能有效维护测试代码、提升软件质量并显著降低维护成本。</span></span></p><p cid="n253" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728930" target="_blank">https://doi.org/10.1145/3728930</a></span></span></p><h3 cid="n254" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">73、Recurring Vulnerability Detection: How Far Are We?</span></span></h3><p cid="n255" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">随着开源软件的快速发展，代码复用已成为加速开发进程的普遍实践。然而这也导致原始漏洞被继承，在复用项目中重现形成复现漏洞（RVs）。传统通用漏洞检测方法受限于可扩展性与适应性，基于学习的方法则常因训练数据集有限而对未见漏洞效果不佳。尽管已有特定复现漏洞检测（RVD）方法被提出，但其针对不同RV特征的有效性尚不明确。本文通过新构建的包含4,569个RVs的大规模数据集（较先前数据集扩展953%）开展实证研究，系统分析RV特征，评估最先进RVD方法的有效性，探究误报/漏报的根本原因并得出关键洞见。基于这些发现，我们设计出新型RVD工具AntMan：通过识别修改函数的显性与隐式调用关系，在函数内实施过程间污点分析和过程内依赖切片以生成综合签名，最终采用柔性匹配检测RVs。评估结果表明该方法具有卓越的有效性、通用性和实用价值。AntMan已检测到4,593个RVs（其中307个获开发者确认），在15个项目中识别出73个新的0-day漏洞，并获得5个CVE标识符。</span></span></p><p cid="n256" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728901" target="_blank">https://doi.org/10.1145/3728901</a></span></span></p><h3 cid="n257" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">74、Reinforcement Learning-Based Fuzz Testing for the Gazebo Robotic Simulator</span></span></h3><p cid="n258" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">作为机器人技术领域应用最广泛的模拟器，Gazebo在开发和测试机器人系统方面发挥着关键作用。鉴于其对机器人操作安全性和可靠性的重大影响，早期缺陷检测至关重要。然而，由于严格的输入结构和庞大的状态空间带来的挑战，直接对Gazebo应用现有模糊测试方法效果有限。本文提出GzFuzz——首个专为Gazebo设计的模糊测试框架。该框架通过语法感知的可行命令生成机制处理严格输入要求，并采用基于强化学习的命令生成器选择机制高效探索状态空间。通过将两种机制整合在统一框架下，GzFuzz能有效检测Gazebo中的缺陷。大量实验表明，GzFuzz在12小时内平均检测到9.6个独特缺陷，其代码覆盖率较现有模糊测试工具AFL++和Fuzzotron实现显著提升，增幅约达239%-363%。在不到六个月的时间内，GzFuzz共发现Gazebo中25个独特崩溃案例，其中24个已获修复或确认。我们的研究成果凸显了直接对Gazebo进行模糊测试的重要性，为此提出了一种新颖高效的方法论，为增强更广泛模拟器的测试能力提供了重要启示。</span></span></p><p cid="n259" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728942" target="_blank">https://doi.org/10.1145/3728942</a></span></span></p><h3 cid="n260" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">75、Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective</span></span></h3><p cid="n261" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">现代软件系统通常具有高度可配置性，以满足不同利益相关者的多样化需求。理解配置项与期望性能属性之间的映射关系，对于提升底层系统的可控性和调优能力具有基础性作用，但由于其黑盒特性，这一直是知识体系中的盲区。尽管已有研究对这些系统进行性能分析，但它们将配置项作为孤立数据点进行分析，未能考虑其固有的空间关联性。这导致无法探查配置空间的许多重要特征，如局部最优区域。本研究提出一种创新视角——将配置空间建模为结构化的&#34;地形景观&#34;。为验证这一理念，我们采用GraphFLA这一基于图数据挖掘的适应度景观分析开源框架，通过对3个真实系统32个运行工作负载中8600万条基准配置进行分析，得出6项主要发现。这些发现共同构建了景观地形的整体图谱，对配置调优和性能建模均具有重要启示意义。</span></span></p><p cid="n262" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728954" target="_blank">https://doi.org/10.1145/3728954</a></span></span></p><h3 cid="n263" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">76、Robust Vulnerability Detection across Compilations: LLVM-IR vs. Assembly with Transformer Model</span></span></h3><p cid="n264" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">检测二进制文件中的漏洞是网络安全领域的一项挑战性任务，尤其在源代码不可用且编译过程及其参数未知的情况下更为困难。现有的基于深度学习的检测方法通常依赖于已知二进制文件的特定编译设置，这可能限制其在其他类型二进制文件上的性能表现。本研究对汇编表示与LLVM-IR进行了全面比较，以确定在编译参数未知时哪种表示更具鲁棒性和适用性。表示方式的选择显著影响检测准确性。本文的另一贡献是采用基于Transformer的模型CodeBERT作为分类工具，用于在编译过程未知的场景下检测漏洞。该研究将Transformer模型应用于LLVM-IR领域中的多类漏洞检测任务，重点关注二进制衍生表示。虽然近期研究已探索了Transformer在源代码和原始二进制指令流漏洞分析中的应用，但作为LLVM-IR层级分类器的系统性评估仍较为有限。先前研究通常依赖基于RNN的方法（此类方法被视为该任务的当前最优方案），但这些模型难以有效捕获长距离依赖关系。为解决这一局限性，我们将基于Transformer的分类方法扩展至二进制文件生成的LLVM-IR，并在此场景下提供全面评估。实验结果凸显了该方法在强化多样化二进制配置系统安全方面的潜力。</span></span></p><p cid="n265" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728903" target="_blank">https://doi.org/10.1145/3728903</a></span></span></p><h3 cid="n266" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">77、RouthSearch: Inferring PID Parameter Specification for Flight Control Program by Coordinate Search</span></span></h3><p cid="n267" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">飞行控制程序被广泛应用于无人机（UAV）中，用于动态管理和维持无人机的飞行行为。这些飞行控制程序包含一个PID控制模块，该模块接收三个用户可配置的PID参数：比例（P）、积分（I）和微分（D）。用户还可在飞行过程中调整这些PID参数以适应不同飞行任务的需求。然而，飞行控制程序对用户提供的PID参数缺乏充分的安全检查，导致无人机存在严重漏洞——输入验证缺陷。当用户错误配置PID参数时，会导致无人机进入危险状态，例如偏离预期路径、失控甚至坠毁。</span></span></p><p cid="n268" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">现有研究通常采用模糊测试等随机测试方法从用户输入中识别无效PID参数。但这些方法在三维PID参数搜索空间中效果有限，且每次无人机测试的动态执行成本极高，进一步影响了随机测试的性能。</span></span></p><p cid="n269" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本研究通过将劳斯-赫尔维茨稳定性判据与坐标搜索相结合，提出名为RouthSearch的方法来解决PID参数错误配置问题。RouthSearch并非以临时方式识别错误配置的PID参数，而是基于原理确定三维PID参数的有效范围。我们首先利用劳斯-赫尔维茨判据识别理论上的PID参数边界，随后通过高效坐标搜索对边界进行精细化处理。RouthSearch确定的三维PID参数有效范围可在飞行过程中过滤用户的错误配置参数，并进一步帮助发现主流飞行控制程序中的逻辑缺陷。</span></span></p><p cid="n270" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们在PX4和ArduPilot两款主流飞行控制程序的八种飞行模式下对RouthSearch进行评估。结果显示：与真实值相比，RouthSearch确定三维PID参数有效范围的准确率达到92.0%。在错误配置参数总数方面，RouthSearch在48小时内发现3,853组PID错误配置，而当前最先进的PGFuzz仅发现449组，性能提升达8.58倍。此外，本方法还帮助检测出ArduPilot和PX4中的三个缺陷。</span></span></p><p cid="n271" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728904" target="_blank">https://doi.org/10.1145/3728904</a></span></span></p><h3 cid="n272" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">78、S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models</span></span></h3><p cid="n273" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">生成式大语言模型（LLM）以其变革性和涌现能力彻底改变了自然语言处理领域。然而，最新研究表明，LLM可能生成违反社会规范的有害内容，这引发了关于部署此类先进模型的安全性与伦理影响的重大关切。因此，在部署前对LLM进行严格全面的安全评估既至关重要又势在必行。尽管存在这一需求，但由于LLM生成空间的广泛性，目前仍缺乏统一规范的风险分类体系来系统反映LLM内容安全性，以及高效探索潜在风险的自动化安全评估技术。</span></span></p><p cid="n274" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为弥补这一显著空白，我们提出S-Eval——一个基于LLM的新型自动化安全评估框架，其配备新定义的全面风险分类体系。S-Eval包含两个核心组件：专家测试LLM Mt和新型安全评判LLM Mc。专家测试LLM Mt负责根据提出的风险管理体系（包含8个风险维度和102项细分风险）自动生成测试用例；安全评判LLM Mc则可提供可量化的可解释安全评估，以增强对LLM风险的认知。与现有工作相比，S-Eval具有三大显著优势：（i）高效性——通过Mt构建包含102类风险共22万个测试用例的多维度开放式基准，并借助Mc对21个具有影响力的LLM进行安全评估，全过程无需人工干预；（ii）有效性——大量验证表明S-Eval能实现更全面的评估和更好的风险感知，Mc不仅能精准量化LLM风险，还提供超越LLaMA-Guard-2等可比模型的可解释深度安全洞察；（iii）适应性——基于LLM的架构使S-Eval可灵活配置，适应LLM快速演进伴随的新安全威胁、测试生成方法和安全评判方法。我们进一步探究超参数和语言环境对模型安全的影响，为未来研究指明方向。目前S-Eval已在工业合作伙伴中部署，为服务数百万用户的多类LLM提供自动化安全评估，实证了其在真实场景中的有效性。</span></span></p><p cid="n275" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728971" target="_blank">https://doi.org/10.1145/3728971</a></span></span></p><h3 cid="n276" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">79、STRUT: Structured Seed Case Guided Unit Test Generation for C Programs using LLMs</span></span></h3><p cid="n277" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">单元测试在缺陷检测与保障软件正确性方面发挥着关键作用，其能帮助开发者在早期发现错误，从而减少软件缺陷。近年来，大型语言模型（LLMs）在自动化单元测试生成领域展现出巨大潜力，但应用LLMs生成单元测试仍面临诸多挑战：1）LLMs生成的测试用例执行通过率较低；2）测试用例覆盖度不足，难以检测代码中的潜在风险；3）现有研究方法主要集中于Java和Python等语言，而对现实世界中至关重要的C语言研究却十分匮乏。为应对这些挑战，我们提出了一种新颖的单元测试生成方法STRUT。该方法以结构化测试用例作为复杂编程语言与LLMs之间的桥梁，通过引导LLMs生成结构化测试用例而非直接生成测试代码，有效缓解了LLMs在生成具有复杂特性编程语言代码时的局限性。具体而言，STRUT首先分析目标方法的上下文并构建结构化种子测试用例，随后引导LLMs生成一组结构化测试用例，最终采用基于规则的方法将结构化测试用例转换为可执行测试代码。通过全面评估，STRUT实现了96.01%的执行通过率、77.67%的代码行覆盖率和63.60%的分支覆盖率，其性能显著优于基于LLMs的基线方法和符号执行工具SunwiseAUnit。这些结果表明STRUT通过融合LLMs优势并克服其固有局限性，具备生成高质量单元测试用例的卓越能力。</span></span></p><p cid="n278" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728970" target="_blank">https://doi.org/10.1145/3728970</a></span></span></p><h3 cid="n279" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">80、SWE-GPT: A Process-Centric Language Model for Automated Software Improvement</span></span></h3><p cid="n280" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大型语言模型（LLM）在代码生成方面展现出卓越性能，显著提升了开发者的编码效率。基于LLM的智能体技术最新进展，推动了端到端自动化软件工程（ASE）的重大突破，特别是在软件维护（如修复缺陷）和演进（如添加新功能）领域。尽管取得这些令人鼓舞的进展，当前研究仍面临两大挑战：其一，最先进性能主要依赖GPT-4等闭源模型，极大限制了技术可及性及在多样化软件工程任务中的定制潜力，同时处理敏感代码库时也引发数据隐私担忧；其二，现有模型主要基于静态代码数据训练，缺乏对软件开发中动态交互、迭代问题解决过程及演进特性的深度理解，导致其在处理复杂项目结构和生成上下文相关解决方案时存在局限，影响实际应用效果。</span></span></p><p cid="n281" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">针对这些挑战，本研究从软件工程视角出发，认识到真实世界的软件维护与演进过程不仅包含静态代码数据，还涉及开发者的思维过程、外部工具使用以及不同职能人员间的交互。我们的目标是开发专为软件改进优化的开源大语言模型，在实现与闭源模型相当性能的同时，提供更强的可访问性和定制潜力。为此，我们推出Lingma SWE-GPT系列模型（包括70亿参数的Lingma SWE-GPT 7B和720亿参数的Lingma SWE-GPT 72B）。通过学习和模拟真实代码提交活动，该系列系统性地融入了软件开发过程中的动态交互与迭代问题解决机制（如仓库理解、故障定位和补丁生成），从而实现对软件改进过程的更全面认知。</span></span></p><p cid="n282" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">基于OpenAI最新提出的SWE-bench-Verified基准测试（包含500个真实GitHub问题），实验结果表明：Lingma SWE-GPT 72B成功解决30.20%的GitHub问题，在自动问题解决方面实现显著提升（较Llama 3.1 405B相对提升22.76%），接近闭源模型性能（GPT-4o解决率为31.80%）；值得注意的是，Lingma SWE-GPT 7B解决率达18.20%，超越Llama 3.1 70B的17.20%，彰显了较小模型在ASE任务中的应用潜力。</span></span></p><p cid="n283" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728981" target="_blank">https://doi.org/10.1145/3728981</a></span></span></p><h3 cid="n284" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">81、Safe4U: Identifying Unsound Safe Encapsulations of Unsafe Calls in Rust using LLMs</span></span></h3><p cid="n285" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">Rust是一种新兴的编程语言，通过严格的编译时检查确保安全性。标记为unsafe的函数表明其具有额外的安全要求（例如已初始化、非空等），社区中称之为契约。这些unsafe函数只能在显式的unsafe代码块中被调用，且契约必须由调用方保证。为了重用并减少unsafe代码，社区实践中推荐采用对unsafe调用的安全封装（EUC）。但若任何契约未得到保证，EUC就会出现缺陷（unsound），可能导致安全Rust中的未定义行为，从而破坏Rust的安全承诺。由于代码与自然语言跨语言理解的局限性，传统技术难以有效识别缺陷EUC。大型语言模型（LLM）虽展现出强大能力，但因契约复杂性及领域知识缺乏，其表现仍不尽如人意。为此，我们提出新型框架Safe4U，融合LLM、静态分析工具与领域知识来识别缺陷EUC。Safe4U首先利用静态分析工具获取相关上下文，随后将原始契约描述分解为多个细粒度分类契约，最终引入领域知识并调用LLM的推理能力验证每个细粒度契约。评估结果表明，Safe4U实现了整体性能提升，且细粒度结果对定位具体缺陷源具有建设性。在真实场景中，Safe4U从CVE报告的11个缺陷EUC中成功识别出9个，并在下载量最高的crates中检测到22个新的缺陷EUC，其中16个已获确认。</span></span></p><p cid="n286" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728890" target="_blank">https://doi.org/10.1145/3728890</a></span></span></p><h3 cid="n287" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">82、Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection</span></span></h3><p cid="n288" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">智能合约漏洞检测是快速发展的区块链领域中的关键挑战。现有漏洞检测方法面临两个主要问题：(1) 现有数据集缺乏全面性和足够质量，漏洞类型覆盖范围有限，且对偏好学习的高质量与低质量解释区分不足；(2) 大语言模型（LLM）往往难以准确解释智能合约安全中的特定概念。通过实证分析，我们发现即使经过持续预训练和监督微调，LLM在精确理解智能合约状态变更执行顺序方面仍存在局限，这可能导致在做出正确检测决策的同时产生错误的漏洞解释。这些局限导致检测性能不佳，进而引发潜在的严重财务损失。  </span></span></p><p cid="n289" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为解决这些挑战，我们提出基于LLaMA-3.1-8B的先进检测方法Smart-LLaMA-DPO。首先，我们构建了涵盖四种漏洞类型及机器不可审计漏洞的综合数据集，包含用于监督微调（SFT）的标签、详细解释和精确漏洞位置，以及用于直接偏好优化（DPO）的配对高质量与低质量输出。其次，我们使用大规模智能合约代码进行持续预训练，以增强LLM对智能合约特定安全实践的理解。进一步，我们利用综合数据集实施监督微调。最后，我们应用DPO技术，通过人类反馈提升生成解释的质量。Smart-LLaMA-DPO采用特殊设计的损失函数，促使LLM增加偏好输出的概率同时降低非偏好输出的概率，从而提升其生成高质量解释的能力。  </span></span></p><p cid="n290" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们在四大漏洞类型（重入、时间戳依赖、整数溢出/下溢和delegatecall）及机器不可审计漏洞上评估Smart-LLaMA-DPO。我们的方法显著优于现有最优基线，F1分数平均提升10.43%，准确率平均提高7.87%。此外，LLM评估与人工评估均证明Smart-LLaMA-DPO生成的解释在正确性、全面性和清晰度方面具有卓越质量。</span></span></p><p cid="n291" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728878" target="_blank">https://doi.org/10.1145/3728878</a></span></span></p><h3 cid="n292" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">83、SoK: A Taxonomic Analysis of DeFi Rug Pulls: Types, Dataset, and Tool Assessment</span></span></h3><p cid="n293" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">拉地毯骗局是去中心化金融（DeFi）领域的关键威胁，造成重大财务损失并侵蚀生态系统信任。尽管研究取得进展，但零散的分类法、有限的数据集和不充分的工具评估仍阻碍着有效检测。通过对学术和行业资源的系统分析，我们建立了包含35种独特拉地毯类型的综合分类法，其中包括9种先前未记录的变体。分析揭示了显著的检测缺口：现有数据集仅覆盖20%的已知类型，这促使我们创建包含2,391个实例的增强数据集，将覆盖率提升至82.9%。对13种检测工具的评估显示其能力存在显著差异（25.7%至62.9%），其中9种类型完全无法检测。最关键的是，面对复杂攻击时工具性能显著下降：单向量攻击的检测率从55.6%骤降至复合场景下的31.3%。这些发现为开发更强大的去中心化系统智能合约漏洞安全测试方法提供了重要见解。</span></span></p><p cid="n294" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728900" target="_blank">https://doi.org/10.1145/3728900</a></span></span></p><h3 cid="n295" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">84、Static Program Reduction via Type-Directed Slicing</span></span></h3><p cid="n296" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">传统的程序切片工具能够针对目标变量构建一个计算相同结果的精简程序变体，即程序切片保留了原始程序的运行时语义。本文提出类型导向切片方法，该方法构建一个更小的程序，确保类型检查器在仅考虑目标程序位置时对切片程序产生相同的结果——即类型导向切片器从特定类型检查器的视角出发，保留了目标程序在编译时的语义。  </span></span></p><p cid="n297" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">类型导向切片是类型检查器设计者和维护者的有效调试辅助工具。当类型检查器在大型代码库上产生意外结果（如崩溃、误报警告、遗漏警告等）时，用户通常会在未提供测试用例的情况下向类型检查器维护者报告错误。当前最先进的程序缩减方案是动态方法：需要反复运行类型检查器以验证最小化结果。而类型导向切片器通过利用类型检查器类型规则固有的模块化特性，无需重新运行类型检查器即可静态解决该问题。我们针对Java开发的类型导向切片原型工具完全自动化，可处理不完整程序，且运行高效。该工具能为三个广泛使用的类型检查器（Java编译器自身、NullAway和Checker Framework）的28个历史缺陷中的25个（89%）生成保留类型检查器异常行为的小型测试用例；在这25个案例中，即使缺少目标程序的类路径，它仍能保持类型检查器的行为特征。此外，在免费层级的CI运行器上，该工具对每个基准测试（代码规模高达数百万行）的处理时间均在一分钟内完成。</span></span></p><p cid="n298" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728968" target="_blank">https://doi.org/10.1145/3728968</a></span></span></p><h3 cid="n299" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">85、Structure-Aware, Diagnosis-Guided ECU Firmware Fuzzing</span></span></h3><p cid="n300" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">电子控制单元（ECU）在现代车辆中扮演着关键角色，其功能范围涵盖基础控制功能至安全关键功能。模糊测试已成为确保ECU固件功能安全与车辆安全性的有效手段。然而现有模糊测试方法主要关注通过外部总线（如CAN）来自其他ECU的输入，却忽视了通过板载总线（如SPI）从内部外设接收的输入。由于输入空间探索受限，这些方法无法全面覆盖ECU固件的模糊测试。此外，现有方法通常缺乏对ECU固件内部状态的可见性，仅依赖有限反馈（如消息超时或硬件指示灯），制约了测试有效性。</span></span></p><p cid="n301" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">针对这些局限性，我们提出结构感知、诊断引导的EcuFuzz框架，以实现全面有效的ECU固件模糊测试。具体而言，EcuFuzz同步处理外部总线（CAN）与板载总线（SPI），利用CAN和SPI的协议结构有效变异CAN消息与SPI序列，并采用基于双核微控制器的外设模拟器处理实时SPI通信。此外，EcuFuzz创新性地引入车辆诊断协议作为反馈机制，通过采集ECU内部状态（包括错误相关变量、故障码及异常上下文）来指导测试进程。在对三家主流一级供应商的十款ECU兼容性评估中，本框架成功适配九款设备；对三款代表性ECU的有效性评估表明，该方法检测出九个未知安全关键故障，相关供应商已发布技术补丁。</span></span></p><p cid="n302" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728914" target="_blank">https://doi.org/10.1145/3728914</a></span></span></p><h3 cid="n303" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">86、Testing the Fault-Tolerance of Multi-sensor Fusion Perception in Autonomous Driving Systems</span></span></h3><p cid="n304" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">生产级自动驾驶系统（如谷歌Waymo与百度Apollo）通常依赖多传感器融合策略实现环境感知。该策略通过结合摄像头与激光雷达的各自优势提升感知鲁棒性，直接影响自动驾驶车辆的安全关键驾驶决策。然而在实际自动驾驶场景中，摄像头与激光雷达均易受各类故障影响，显著改变自动驾驶系统的决策与后续行为。开发阶段需全面测试多传感器融合的鲁棒性。现有测试方法仅关注系统未能识别的极端案例，尚未深入研究传感器故障如何影响自动驾驶系统的整体行为。为此，我们提出FADE——首个全面评估基于多传感器融合感知的自动驾驶系统容错能力的测试方法。我们系统化构建自动驾驶车辆摄像头与激光雷达的故障模型，并将这些故障注入基于多传感器融合的自动驾驶系统，以测试其在多种场景下的行为。为高效探索传感器故障模型的参数空间，我们设计了一种反馈引导的差分模糊测试器，用于揭示注入故障引发的自动驾驶系统安全违规。我们在代表性工业级自动驾驶系统百度Apollo上评估FADE，实验结果证明了该方法的实用价值并揭示重要发现。我们进一步使用百度Apollo 6.0 EDU自动驾驶车辆进行实体实验，在真实场景中验证了这些发现。</span></span></p><p cid="n305" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728910" target="_blank">https://doi.org/10.1145/3728910</a></span></span></p><h3 cid="n306" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">87、The First Prompt Counts the Most! An Evaluation of Large Language Models on Iterative Example-Based Code Generation</span></span></h3><p cid="n307" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大型语言模型（LLM）在代码生成方面的能力已得到广泛研究，尤其是在根据自然语言描述实现目标功能方面。作为自然语言的替代方案，输入输出（I/O）示例提供了一种可访问、明确且灵活的功能描述方式。然而，其固有的多样性、不透明性和不完整性为理解和实现目标需求带来了更大挑战。因此，基于I/O示例生成代码（即基于示例的代码生成）提供了新视角，使我们能够额外评估LLM从有限信息推断目标功能以及处理新型需求的能力。但关于LLM在基于示例代码生成中的相关研究仍处于探索阶段。为填补这一空白，本文首次对基于示例的代码生成开展综合性研究。针对I/O示例不完整性导致的错误问题，我们采用迭代评估框架，并将基于示例的代码生成目标形式化为两个连续子目标：生成符合给定示例的代码，以及通过（迭代）给定示例成功实现目标功能的代码。我们使用包含172个多样化目标功能（源自HumanEval和CodeHunt）的新基准测试评估了六个前沿LLM。结果表明：当使用迭代I/O示例而非自然语言描述需求时，LLM的得分下降超过60%，说明基于示例的代码生成对当前LLM仍具挑战性。值得注意的是，绝大多数（甚至超过95%）成功实现的功能均在首轮迭代中完成，表明LLM难以有效利用迭代补充的需求。此外，我们发现将I/O示例与即使不精确且碎片化的自然语言描述结合可显著提升LLM性能，且初始I/O示例的选择也会影响得分，这为提示优化提供了可能。这些发现凸显了交互过程中早期提示的重要性，并为增强基于LLM的代码生成提供了关键见解与启示。</span></span></p><p cid="n308" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728947" target="_blank">https://doi.org/10.1145/3728947</a></span></span></p><h3 cid="n309" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">88、The Incredible Shrinking Context... in a Decompiler Near You</span></span></h3><p cid="n310" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">二进制代码反编译在以太坊虚拟机（EVM）智能合约领域已成为一项至关重要的应用。出于多种逆向工程或工具开发目的，几乎每年都会出现新的主流反编译器并广受欢迎。从技术角度看，该问题具有根本性挑战：其核心是从高度优化的延续传递风格（CPS）表示中恢复高级控制流。在架构层面，反编译器可通过静态分析或符号执行技术构建。  </span></span></p><p cid="n311" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们提出Shrnkr——一种基于静态分析的反编译器，继承了最先进的Elipmoc反编译器的优势。Shrnkr在所有关键维度上均实现了显著改进：可扩展性、完整性和精确度。其核心技术采用了一种新型静态分析上下文变体：收缩上下文敏感性。该技术通过深度裁剪静态分析上下文，主动&#34;遗忘&#34;控制流历史，从而为更精确的推理创造空间。  </span></span></p><p cid="n312" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们将Shrnkr与基于静态分析和符号执行的最先进反编译器进行对比。在标准基准测试集中，Shrnkr可扩展至99.5%以上的合约（Elipmoc约为95%），代码覆盖率（即触及并成功反编译的代码）较Heimdall-rs提升67%，关键不精确度指标较Elipmoc降低超65%。</span></span></p><p cid="n313" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728935" target="_blank">https://doi.org/10.1145/3728935</a></span></span></p><h3 cid="n314" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">89、Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection</span></span></h3><p cid="n315" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">根据我们对漏洞检测机器学习（ML4VD）的调研，过去五年中发表的论文十有八九将ML4VD定义为函数级二元分类问题：给定一个函数，判断其是否包含安全漏洞？基于安全研究者的经验，在判定某函数是否导致程序存在攻击漏洞时，我们往往需要先理解该函数的调用上下文。本文通过分析主流ML4VD数据集中的漏洞函数与非漏洞函数，探究在缺乏上下文的情况下做出准确判断的实际可行性。若某函数因实际安全漏洞的修复而被修改，且被确认为导致程序漏洞的根源，则判定为漏洞函数；反之则为非漏洞函数。研究发现，几乎所有案例均表明脱离上下文无法做出准确判断：漏洞函数往往仅因存在诱发漏洞的调用上下文而具有危险性，而非漏洞函数在特定上下文中也可能转化为漏洞函数。  </span></span></p><p cid="n316" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">但为何现有ML4VD技术能在样本明显信息不足的情况下获得高评分？伪相关性现象：研究发现即使仅依据词频统计也能获得高评分，这表明当前数据集存在被利用以获取高评分而非真正检测安全漏洞的缺陷。本文结论指出，主流ML4VD问题定义存在根本性缺陷，并质疑该领域大量研究的内部有效性。建设性地，我们呼吁建立更有效的基准评估方法以衡量ML4VD的真实能力，提出替代性问题定义框架，并探讨其对机器学习与程序分析研究评估体系的更广泛启示。</span></span></p><p cid="n317" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728887" target="_blank">https://doi.org/10.1145/3728887</a></span></span></p><h3 cid="n318" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">90、Tracezip: Efficient Distributed Tracing via Trace Compression</span></span></h3><p cid="n319" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">分布式追踪是云服务系统监控与测试的基础构建模块。为降低计算与存储开销，当前普遍采用采样方式减少追踪数据采集。然而现有工作面临追踪完整性与系统开销之间的权衡：基于头部采样的方法在请求进入系统时 indiscriminately 选择追踪对象，可能遗漏关键事件；基于尾部采样的方法先全量采集请求，再选择性保留边缘案例追踪，但会带来追踪数据收集与录入的开销。本文另辟蹊径，提出Tracezip通过追踪压缩提升分布式追踪效率。核心洞见在于：追踪数据间存在显著冗余，导致相同数据在服务与后端之间重复传输。我们设计了一种名为跨度检索树（SRT）的新型数据结构，可在服务端持续封装此类冗余，将追踪跨度转换为轻量形式。在后端，通过检索先前跨度已传输的公共数据即可无缝重构完整追踪链。Tracezip包含一系列优化SRT结构的策略，以及通过差分更新机制高效同步服务与后端间SRT的方法。基于微服务基准测试、主流云服务系统和真实生产追踪数据的评估表明，Tracezip能以可忽略的开销显著提升追踪收集性能。我们已在OpenTelemetry Collector中实现Tracezip，使其与现有追踪API保持兼容。</span></span></p><p cid="n320" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728888" target="_blank">https://doi.org/10.1145/3728888</a></span></span></p><h3 cid="n321" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">91、Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test Oracles</span></span></h3><p cid="n322" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出Tratto——一种神经符号方法，能够从源代码和文档中生成可作为公理化预言（布尔表达式）的断言。Tratto的符号模块利用编程语言的语法、被测单元及其上下文（所属类与可用API）来约束可成功生成有效预言词符的搜索空间。其神经模块采用经微调的Transformer模型，既决策是否输出预言，又从符号模块返回的词符集合中选择下一个词符来逐步构建预言。实验表明，Tratto以73%准确率、72%精确率和61% F1分数显著优于现有公理化预言生成方法，大幅超越本研究中最优符号方法与神经方法的最佳结果（分别为61%、62%和37%）。Tratto生成的公理化预言数量是当前符号方法的三倍，而生成的误报数量比采用少样本学习和思维链提示的GPT-4减少十倍。</span></span></p><p cid="n323" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728960" target="_blank">https://doi.org/10.1145/3728960</a></span></span></p><h3 cid="n324" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">92、Type-Alias Analysis: Enabling LLVM IR with Accurate Types</span></span></h3><p cid="n325" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">LLVM中间表示（IR）作为LLVM编译器基础设施的核心，提供了强大的类型系统和静态单赋值（SSA）形式，非常适合程序分析。但其单类型设计为每个IR变量严格指定单一类型，即使该变量可能合法对应多种类型。近期不透明指针的引入加剧了这一局限：IR中所有指针均以通用指针类型（ptr）统一表示，抹除了具体指针目标类型信息，导致许多基于类型的分析失效。</span></span></p><p cid="n326" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">为突破单类型设计的限制，我们提出类型别名分析——一种多类型设计方案，通过维护IR变量的类型别名集并在IR指令间推断类型。我们开发了原型工具TypeCopilot，专门针对C程序生成的启用不透明指针的LLVM IR恢复具体指针目标类型。TypeCopilot实现了98.57%的准确率和94.98%的覆盖率，使现有分析工具在采用不透明指针后仍能保持有效性。为促进进一步研究和安全应用，我们已开源TypeCopilot，为社区在现代LLVM IR上开展精确的类型感知安全分析提供实践基础。</span></span></p><p cid="n327" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728974" target="_blank">https://doi.org/10.1145/3728974</a></span></span></p><h3 cid="n328" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">93、Uncovering API-Scope Misalignment in the App-in-App Ecosystem</span></span></h3><p cid="n329" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">摘要“应用内应用”（app-in-app）模式是移动系统的新兴趋势，超级应用（简称superApps，如微信、百度、抖音）通过提供特权API，允许外部供应商在其平台上开发小程序（简称miniApps）。为便于管理，超级应用设计了特定的权限配置（称为scope）来授予API对特定功能和资源的访问权限。在API实现过程中严格遵守这些权限范围对维护安全至关重要，否则超级应用的权限管理可能被绕过——我们将这种漏洞称为API-权限范围失配。本研究首次对应用内应用生态中的API-权限范围失配问题进行系统性分析，揭示了根本原因和安全风险。更重要的是，我们开发了名为ScopeChecker的自动化工具，用于检测超级应用和小程序中的API-权限范围失配问题。该工具通过将Android权限机制集成到超级应用功能中，提取标准API-权限范围映射关系，并基于LLM的代码生成技术创建可执行的API代码片段作为测试用例。执行结果反映了API与权限范围的实际映射关系，通过与标准映射比对即可识别失配现象。随后，ScopeChecker通过将失配API与定制化的目标小程序方法导向抽象语法树（MAST）进行匹配，验证小程序中的失配情况。经人工确认，ScopeChecker在头部超级应用中检出38个存在权限失配的API，其性能优于当前最先进的小程序测试方法。值得注意的是，我们获得了超级应用开发商及CNVD的11次正面回应，其中9个漏洞获得确认并获奖励：包括1个高风险、7个中风险和1个低风险漏洞。为评估普遍性，ScopeChecker检测了42,000余个小程序，发现51%存在API-权限范围失配问题，平均每个小程序存在1.4个失配API。最后，我们通过分析真实攻击案例，阐述了由API-权限范围失配引发的四类安全威胁。</span></span></p><p cid="n330" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728962" target="_blank">https://doi.org/10.1145/3728962</a></span></span></p><h3 cid="n331" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">94、Understanding Model Weaknesses: A Path to Strengthening DNN-Based Android Malware Detection</span></span></h3><p cid="n332" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">Android恶意软件检测仍是网络安全研究中的关键挑战。近年来研究利用人工智能技术（特别是深度神经网络DNN）训练检测模型，但常用训练数据集中恶意软件家族间的显著不平衡往往会影响其有效性。这种不平衡导致模型在主流类别上过拟合，而在低代表性类别上表现不佳，增加了对罕见恶意软件家族的预测不确定性。为改善许多DNN模型的次优性能，我们提出MalTutor新型框架，通过优化训练流程增强模型鲁棒性。我们的核心洞见在于将不确定性从&#34;负担&#34;转化为&#34;资产&#34;，并将其策略性融入DNN训练方法。具体而言，我们首先评估DNN模型在不同训练周期中的预测不确定性，以此指导样本分类。结合课程学习策略，我们从低不确定性的易学习样本开始训练，逐步加入高不确定性的难学习样本。实验结果表明，MalTutor显著提升了在不平衡数据集上训练的模型性能：准确率提高31.0%，F1分数提升138.8%，特别是在检测各类恶意应用时的平均准确率提升133.9%。我们的发现为利用不确定性增强面向预测的软件工程任务中DNN模型鲁棒性提供了重要见解。</span></span></p><p cid="n333" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728884" target="_blank">https://doi.org/10.1145/3728884</a></span></span></p><h3 cid="n334" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">95、Understanding Practitioners’ Expectations on Clear Code Review Comments</span></span></h3><p cid="n335" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">代码审查评论（CRC）在现代代码审查过程中至关重要。它为审查者提供了识别潜在缺陷、提供建设性反馈及改进建议的机会。清晰简洁的代码审查评论能促进开发者之间的沟通，并对正确理解已发现问题和建议解决方案具有关键作用。尽管CRC清晰度的重要性已被广泛认可，但目前仍缺乏关于良好清晰度构成要素及评估标准的指导原则。本文通过综合研究来理解和评估CRC的清晰度：首先基于文献综述和实践者调研，推导出与CRC清晰度相关的一组属性——RIE属性（即相关性、信息量和表达质量）及其对应评估标准；随后对九种编程语言开源项目中的CRC清晰度进行实证分析，发现其中较大比例（28.8%）的评论至少存在某一属性上的清晰度缺陷；最后，我们通过提出ClearCRC框架探索自动评估CRC清晰度的可行性。实验结果表明，基于预训练语言模型的ClearCRC能有效评估CRC清晰度，其平衡准确率最高达73.04%，F-1分数最高达94.61%。</span></span></p><p cid="n336" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728931" target="_blank">https://doi.org/10.1145/3728931</a></span></span></p><h3 cid="n337" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">96、Unlocking Low Frequency Syscalls in Kernel Fuzzing with Dependency-Based RAG</span></span></h3><p cid="n338" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">大多数覆盖引导的内核模糊测试工具通过系统调用序列合成来测试操作系统内核。然而，在模糊测试过程中仍存在极少或未被覆盖的系统调用（称为低频系统调用，LFS），这意味着相关代码分支未被探索。这是由于LFS的复杂依赖性和突变不确定性，使得模糊测试工具难以生成相应的系统调用序列。由于许多内核模糊测试工具能够基于选择表机制从当前语料库中动态学习系统调用依赖关系，提供全面且高质量的种子有助于覆盖LFS。但构建此类种子严重依赖专家经验来解决系统调用依赖关系。  </span></span></p><p cid="n339" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">本文提出SyzGPT，首个通过大语言模型（LLM）自动为LFS生成有效种子的内核模糊测试框架。我们采用基于依赖关系的检索增强生成（DRAG）方法释放LLM的潜力，并设计了一系列步骤提升生成种子的有效性。首先，SyzGPT通过LLM从现有文档中自动提取系统调用依赖关系；其次，基于依赖关系从模糊测试语料库中检索程序，为LLM构建自适应上下文；最后，通过反馈周期性生成并修复种子以丰富LFS的模糊测试语料库。我们提出了一套针对内核领域种子生成的新评估指标。实验表明，SyzGPT能生成有效率达87.84%的种子，并可扩展到离线和微调LLM。与七种最先进的内核模糊测试工具相比，SyzGPT平均将代码覆盖率提升17.73%，LFS覆盖率提升58.00%，漏洞检测能力提高323.22%。此外，SyzGPT独立发现了26个未知内核漏洞（其中10个与LFS相关），11个已获确认。</span></span></p><p cid="n340" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728913" target="_blank">https://doi.org/10.1145/3728913</a></span></span></p><h3 cid="n341" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">97、Validating Network Protocol Parsers with Traceable RFC Document Interpretation</span></span></h3><p cid="n342" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">验证网络协议实现的正确性极具挑战性，主要存在预测基准缺失（oracle）和可追溯性两大难题。前者决定了何时应将协议实现判定为存在缺陷——尤其当错误未引发任何可观测症状时；后者则帮助开发者理解实现如何违反协议规范，从而促进错误修复。与现有研究很少同时考虑这两个问题不同，本文基于大语言模型（LLM）的最新进展，同时解决这两个问题并提供有效方案。我们的核心发现是：网络协议通常随结构化规范文档（即RFC文档）发布，这些文档可通过LLM系统性地转换为形式化的协议消息规范。此类规范虽可能因LLM幻觉存在误差，但可作为准预测基准来验证协议解析器，而验证结果又会逐步优化该基准。由于基准源自规范文档，我们在协议实现中发现的任何错误均可追溯至文档，从而解决可追溯性问题。我们使用九种网络协议及其C、Python和Go语言实现进行了广泛评估。结果表明：本方法优于现有最优技术，共检测出69个错误（其中36个已确认）。本项目还展示了基于自然语言规范实现全自动化软件验证的潜力——该过程因需理解规范文档并推导测试输入的预期输出，历来被认为主要依赖人工完成。</span></span></p><p cid="n343" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728955" target="_blank">https://doi.org/10.1145/3728955</a></span></span></p><h3 cid="n344" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">98、VerLog: Enhancing Release Note Generation for Android Apps using Large Language Models</span></span></h3><p cid="n345" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">发布说明是向用户和开发者传达软件更新细节的重要文档，但其生成过程仍耗时且易出错。本文提出VerLog，一种利用大语言模型（LLM）增强软件发布说明生成的新技术。VerLog通过自适应提示的少样本上下文学习，激发LLM的图推理能力，使其能准确解读并记录代码变更的语义信息。此外，VerLog融合了多粒度信息（包括细粒度代码修改和高层非代码工件）以指导生成过程，确保发布说明具备全面性、准确性和可读性。我们将VerLog应用于248个独特Android应用的42个版本，并进行了广泛评估。结果表明，无论是在高质量参考发布说明的受控实验中，还是野外评估中，VerLog在生成发布说明的完整性、准确性、可读性及整体质量上均显著优于现有基线方法（精确率、召回率和F1值最高提升18%–21%）。</span></span></p><p cid="n346" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728961" target="_blank">https://doi.org/10.1145/3728961</a></span></span></p><h3 cid="n347" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">99、Walls Have Ears: Demystifying Notification Listener Usage in Android Apps</span></span></h3><p cid="n348" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">Android系统中的通知监听服务（NLS）允许第三方应用监控和处理设备通知，虽能实现强大功能，但也带来安全与隐私风险。尽管访问NLS需特殊权限，其仍被恶意行为者反复利用。然而，目前缺乏对NLS使用模式及其安全影响的系统性研究。本文提出NLRadar——一种结合静态分析与大语言模型（LLM）的混合方法，用于检测Android应用中的NLS使用情况。我们将NLRadar应用于大规模应用（含恶意软件与常规应用），以揭示NLS使用模式并挖掘滥用行为。分析表明NLS存在严重滥用现象，包括应用不安全存储社交媒体消息、利用NLS进行破坏性竞争或窃取短信凭证，以及通过NLS传播推广信息甚至恶意链接等。研究还发现应用更新中存在未公开的NLS使用变更，且隐私政策披露不充分。这些发现表明，亟需加强对NLS使用的严格审查，并提升开发者对负责任NLS实践的认知。</span></span></p><p cid="n349" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728898" target="_blank">https://doi.org/10.1145/3728898</a></span></span></p><h3 cid="n350" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">100、Wemby’s Web: Hunting for Memory Corruption in WebAssembly</span></span></h3><p cid="n351" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">WebAssembly通过原生代码实现了Web应用中性能关键模块的高速执行。然而最新研究表明，WebAssembly模块中的内存破坏错误可能被用于攻击Web应用。本文首次对WebAssembly内存破坏问题展开系统性分析，揭示了一种新型威胁模型的普遍存在：攻击者可通过内存破坏实现受害者浏览器端的代码注入。通过对37,797个域名的大规模分析，我们发现有29,411个（77.81%）域名完全信任来自潜在攻击者控制源的数据。攻击者可利用内存错误操纵WebAssembly内存——这些被隐式信任的数据常被传入敏感函数（如eval）或通过innerHTML直接插入DOM。因此，攻击者可滥用这种信任实现JavaScript代码执行，即跨站脚本攻击（XSS）。</span></span></p><p cid="n352" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">针对该问题，我们提出首个整体分析WebAssembly网站的有效方案Wemby。通过模糊测试技术，Wemby能有效检测Web应用中远程暴露的内存破坏错误。我们实现了无需源码的WebAssembly插桩方案，提供细粒度的内存破坏检测能力。在实际应用中，Wemby成功发现了多个内存破坏漏洞（包括Zoom平台漏洞）。性能评估表明：Wemby相比现有WebAssembly模糊测试工具有显著提升，平均速度提高232倍，代码覆盖率增加46%。</span></span></p><p cid="n353" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728937" target="_blank">https://doi.org/10.1145/3728937</a></span></span></p><h3 cid="n354" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">101、What Happened in This Pipeline? Diffing Build Logs with CiDiff</span></span></h3><p cid="n355" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">持续集成（CI）被开发者广泛用于确保软件项目的质量与可靠性。然而，诊断CI回归是一个繁琐的过程，需要人工分析冗长的构建日志。本文探索了文本差异分析如何辅助CI回归调试。由于现成的差异比对算法效果欠佳，我们提出了一种专为构建日志设计的差异化算法CiDiff。我们在包含17,906个CI回归案例的新数据集上，通过准确率研究、量化分析和用户调研，将CiDiff与多种基线方法进行对比。结果表明：在中等规模案例中，我们的算法可将需要检查的代码行数减少约60%，与当前主流的LCS差异算法相比具有合理开销。最终，在70%的回归案例中大多数参与者倾向于选择我们的算法，而LCS差异算法仅获得5%的偏好选择。</span></span></p><p cid="n356" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728966" target="_blank">https://doi.org/10.1145/3728966</a></span></span></p><h3 cid="n357" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">102、Why Does My Transaction Fail? A First Look at Failed Transactions on the Solana Blockchain</span></span></h3><p cid="n358" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">Solana是一个新兴的区块链平台，以其高吞吐量和低交易成本著称，已成为去中心化金融（DeFi）、非同质化代币（NFT）及其他Web 3.0应用的首选基础设施。在Solana生态中，交易发起者通过提交多样化指令与各类智能合约交互，其中包括采用自动化做市商（AMM）机制的去中心化交易所（DEX），使用户无需中介即可直接在链上交易加密货币。尽管Solana具备高吞吐量和低成本优势，这些特性却使其面临机器人滥发交易以牟利的问题，导致交易失败和网络拥堵现象频发。  </span></span></p><p cid="n359" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">现有研究主要集中于Solana区块链的性能评估（特别是可扩展性与交易吞吐量）以及智能合约安全性改进，而对失败交易的特征及影响尚缺乏深入探索。为此，我们基于涵盖7200万个区块中超过15亿笔失败交易的精选数据集，对Solana失败交易展开大规模实证研究。具体而言，我们首先从交易发起者、触发失败的程序和时间模式三个维度刻画失败交易特征，并将其与成功交易的区块位置和交易成本进行对比；随后根据错误日志中的报错信息对失败交易分类，并探究特定程序与交易发起者如何关联这些错误。  </span></span></p><p cid="n360" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们发现：Solana交易失败率呈现日周期性波动，且与失败交易量呈强正相关，其中机器人交易失败率高达58.43%；失败交易错误日志中存在十类典型错误，因&#34;价格/利润未满足&#34;和&#34;无效状态&#34;导致的失败占比达67.18%；AMM在失败交易中主要遭遇&#34;无效状态&#34;错误，而DEX聚合器更易受&#34;价格/利润未满足&#34;错误影响；交易发起者中，机器人因高频交易和复杂合约交互面临更广泛的错误类型，普通用户则错误类型较为有限。基于研究结果，我们提出降低Solana交易失败率的实践建议，并展望未来研究方向。</span></span></p><p cid="n361" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728943" target="_blank">https://doi.org/10.1145/3728943</a></span></span></p><h3 cid="n362" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">103、WildSync: Automated Fuzzing Harness Synthesis via Wild API Usage Recovery</span></span></h3><p cid="n363" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">模糊测试是高效测试软件的最实用技术之一。当对软件库API进行模糊测试时，高质量的测试套件至关重要，它能使模糊器以精确的调用序列和函数参数执行API。尽管开发人员通常依赖人工编写测试套件，但自动化生成方法正受到日益广泛的关注。现有研究因依赖基于编译器的分析或运行时执行轨迹（需要人工设置配置），在可扩展性和有效性方面存在局限。我们对多个活跃测试库的研究表明，大量被开源项目实际使用的导出API函数尚未被现有测试套件或单元测试覆盖。这些API函数缺乏测试会增加漏洞未被发现的风险，进而可能引发安全问题。为改善现有模糊测试方法的覆盖不足问题，我们提出一种创新方法：通过从真实应用场景中提取未测试函数的使用模式，基于轻量级抽象语法树分析技术从外部源代码中提取API使用规范，并将这些使用模式集成到现有测试套件中以构建覆盖未测试函数的新套件。我们实现了名为WildSync的原型系统，能够为OSS-Fuzz上的C/C++库自动生成测试套件。实验表明，WildSync成功为OSS-Fuzz中24个活跃测试库及3个可后续集成的主流库生成469个新测试套件，覆盖函数数量增加超过1.3千个，代码行数增加超过1.6万行，同时发现了7个此前未检测到的漏洞。</span></span></p><p cid="n364" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728918" target="_blank">https://doi.org/10.1145/3728918</a></span></span></p><h3 cid="n365" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">104、You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects</span></span></h3><p cid="n366" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">执行项目测试套件的能力在许多场景中至关重要，例如评估代码质量与覆盖率、验证开发者或自动化工具提交的代码变更、确保与依赖项的兼容性等。尽管其重要性显著，但在实践中执行项目测试套件常面临挑战，因为不同项目采用不同的编程语言、软件生态、构建系统、测试框架及其他工具。这些挑战使得创建一种适用于不同项目的可靠通用测试执行方法变得困难。本文提出ExecutionAgent，这是一种自动化技术，能够通过源代码为任意项目构建测试脚本并运行其测试用例。受人类开发者处理该任务方式的启发，我们的方法基于大型语言模型（LLM）构建自主代理，可自动执行命令并与主机系统交互。该代理通过元提示（meta-prompting）技术获取与目标项目相关的最新技术指南，并基于前序步骤的反馈迭代优化其执行流程。我们在评估中将ExecutionAgent应用于50个开源项目，这些项目涵盖14种编程语言及多种构建与测试工具。该方法成功执行了33/50项目的测试套件，且与基准测试套件执行结果的偏差仅为7.5%。相比现有最佳技术，该成果将成功率提升了6.6倍。该方法成本可控，单项目平均执行时间为74分钟，LLM调用成本仅为0.16美元。我们期望ExecutionAgent能成为开发者、自动化编程工具及研究人员的实用工具，助力跨多样项目的测试执行需求。</span></span></p><p cid="n367" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728922" target="_blank">https://doi.org/10.1145/3728922</a></span></span></p><h3 cid="n368" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">105、ZTaint-Havoc: From Havoc Mode to Zero-Execution Fuzzing-Driven Taint Inference</span></span></h3><p cid="n369" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="softbreak" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "></span><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">模糊测试是一种流行的软件漏洞发现技术，其核心问题在于识别能影响程序行为的关键字节。污点分析能以白盒方式追踪关键字节的数据流，但常存在稳定性问题且无法在大型现实程序中运行。模糊驱动污点推断（FTI）是一种简单的黑盒技术，通过监控程序执行实例的动态行为，以黑盒方式推断关键字节。然而该方法需要额外O(N)次程序执行，导致较大运行时开销。  </span></span></p><p cid="n370" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">我们观察到模糊测试中广泛使用的突变方案——havoc模式，可转化为零额外执行开销的轻量级FTI。本研究首先提出havoc模式的计算模型，形式化描述其突变过程。基于该模型，我们证明havoc模式能在生成和执行新测试用例的同时启动FTI，进而提出无需额外程序执行的ZTaint-Havoc新型FTI方案。ZTaint-Havoc在UniBench和FuzzBench上的插装开销仅分别为3.84%和12.58%。最后我们提出基于ZTaint-Havoc识别关键字节的高效突变算法。  </span></span></p><p cid="n371" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">通过综合评估havoc模式的计算模型，我们验证了将其转化为零额外执行开销的高效FTI的可行性。基于AFL++的havoc模式实现原型ZTaint-Havoc，并在FuzzBench和UniBench数据集上进行评估。大量实验结果表明：在24小时测试中，ZTaint-Havoc相较原生AFL++在FuzzBench和UniBench上的边覆盖率最高提升33.71%和51.12%，平均提升分别为2.97%和6.12%。</span></span></p><p cid="n372" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728916" target="_blank">https://doi.org/10.1145/3728916</a></span></span></p><h3 cid="n373" mdtype="heading" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">106、xFUZZ: A Flexible Framework for Fine-Grained, Runtime-Adaptive Fuzzing Strategy Composition</span></span></h3><p cid="n374" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">模糊测试是检测软件漏洞最高效的技术之一。现有方法在不同目标间存在性能不一致问题，且依赖于僵化的粗粒度模糊测试策略组合，限制了在运行时自适应融合不同模糊测试策略优势的灵活性。为解决这些挑战，我们提出了一个支持细粒度运行时自适应策略组合的灵活可扩展模糊测试框架。该框架将主流输入调度与变异调度策略集成为可独立切换的细粒度插件，使用户能在整个测试过程中自适应替换任意插件。此外，我们提出基于滑动窗口汤普森采样的自适应算法，在测试过程中动态选择最优的模糊测试策略组合。实验结果表明：该框架在独特漏洞发现数量上较最先进模糊测试工具提升10.07%，代码覆盖率提高4.94%。值得注意的是，其在测试套件37个漏洞中率先检测出21个，证明了其在多样化目标上的有效性。</span></span></p><p cid="n375" mdtype="paragraph" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span md-inline="plain" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf="">链接：</span></span><span md-inline="url" spellcheck="false" style=" box-sizing: border-box; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; "><span leaf=""><a href="https://doi.org/10.1145/3728873" target="_blank">https://doi.org/10.1145/3728873</a></span></span></p><p style="display: none;"><mp-style-type data-value="3"></mp-style-type></p>



<p><a href="2247486003">阅读原文</a></p>
<p><a href="https://wechat2rss.xlab.app/link-proxy/?k=fb645cec&amp;r=1&amp;u=https%3A%2F%2Fmp.weixin.qq.com%2Fs%3F__biz%3DMzU0MzgzNTU0Mw%3D%3D%26mid%3D2247486003%26idx%3D1%26sn%3D82d1280ff69952f09d94eb5f9ff2d59a">跳转微信打开</a></p>
]]></content:encoded>
      <pubDate>Sat, 20 Sep 2025 21:58:00 +0800</pubDate>
    </item>
  </channel>
</rss>