<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Yummytanmo</title>
        <link>http://preview.tangly1024.com/</link>
        <description>Learning and wandering ~</description>
        <lastBuildDate>Wed, 11 Jun 2025 03:51:28 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en-US</language>
        <copyright>All rights reserved 2025, Wenxuan Wang</copyright>
        <item>
            <title><![CDATA[Efficient Long CoT Reasoning in Small Language Models]]></title>
            <link>http://preview.tangly1024.com/article/20ed0968-b245-808c-98c5-d3db5fbbcfc7</link>
            <guid>http://preview.tangly1024.com/article/20ed0968-b245-808c-98c5-d3db5fbbcfc7</guid>
            <pubDate>Tue, 10 Jun 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-20ed0968b245808c98c5d3db5fbbcfc7"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-20ed0968b24580a2bdc2fcd8402c5804" data-id="20ed0968b24580a2bdc2fcd8402c5804"><span><div id="20ed0968b24580a2bdc2fcd8402c5804" class="notion-header-anchor"></div><a class="notion-hash-link" href="#20ed0968b24580a2bdc2fcd8402c5804" title="Open AI o1, QWQ  和 Deepseek-R1 scale up the length of CoT steps 显著提高了推理表现"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Open AI o1, QWQ  和 Deepseek-R1 scale up the length of CoT steps 显著提高了推理表现</span></span></h3><div class="notion-row notion-block-20ed0968b2458034b024d25f5381f29e"><div class="notion-column notion-block-20ed0968b2458032bf86c59a0dffe0fb" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b2458020af9fd28912aaf543" href="https://arxiv.org/abs/2201.11903"><div><div class="notion-bookmark-title">Chain-of-Thought Prompting Elicits Reasoning in Large Language Models</div><div class="notion-bookmark-description">We explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning. In particular,...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-8020-af9f-d28912aaf543&amp;t=20ed0968-b245-8020-af9f-d28912aaf543" alt="Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2201.11903</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-8020-af9f-d28912aaf543&amp;t=20ed0968-b245-8020-af9f-d28912aaf543" alt="Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div><div class="notion-column notion-block-20ed0968b24580d791ffd6c092af6c95" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b245801699f0c78db21fe142" href="https://arxiv.org/abs/2203.11171"><div><div class="notion-bookmark-title">Self-Consistency Improves Chain of Thought Reasoning in Language Models</div><div class="notion-bookmark-description">Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy,...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-8016-99f0-c78db21fe142&amp;t=20ed0968-b245-8016-99f0-c78db21fe142" alt="Self-Consistency Improves Chain of Thought Reasoning in Language Models" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2203.11171</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-8016-99f0-c78db21fe142&amp;t=20ed0968-b245-8016-99f0-c78db21fe142" alt="Self-Consistency Improves Chain of Thought Reasoning in Language Models" loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div></div><div class="notion-text notion-block-20ed0968b245807db595e41e5243beca">(Wei et al., 2022b; Wang et al., 2023a; Kojima et al., 2022) CoT prompting</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-20ed0968b245805abdffe9940aa3bb2f" data-id="20ed0968b245805abdffe9940aa3bb2f"><span><div id="20ed0968b245805abdffe9940aa3bb2f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#20ed0968b245805abdffe9940aa3bb2f" title="对SLM提出新的Challenge"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">对SLM提出新的Challenge</span></span></h3><div class="notion-text notion-block-20ed0968b2458080b417d812f318ec2c">They also introduce new challenges to small language models (SLMs) with about 7B parameters which often use distillation methods to learn such long CoT reasoning (Guo et al., 2025; Face, 2025).</div><a target="_blank" rel="noopener noreferrer" href="https://github.com/huggingface/open-r1" class="notion-external notion-external-block notion-row notion-block-20ed0968b2458071b7cef27edb471e4f"><div class="notion-external-image"><svg viewBox="0 0 260 260"><g><path d="M128.00106,0 C57.3172926,0 0,57.3066942 0,128.00106 C0,184.555281 36.6761997,232.535542 87.534937,249.460899 C93.9320223,250.645779 96.280588,246.684165 96.280588,243.303333 C96.280588,240.251045 96.1618878,230.167899 96.106777,219.472176 C60.4967585,227.215235 52.9826207,204.369712 52.9826207,204.369712 C47.1599584,189.574598 38.770408,185.640538 38.770408,185.640538 C27.1568785,177.696113 39.6458206,177.859325 39.6458206,177.859325 C52.4993419,178.762293 59.267365,191.04987 59.267365,191.04987 C70.6837675,210.618423 89.2115753,204.961093 96.5158685,201.690482 C97.6647155,193.417512 100.981959,187.77078 104.642583,184.574357 C76.211799,181.33766 46.324819,170.362144 46.324819,121.315702 C46.324819,107.340889 51.3250588,95.9223682 59.5132437,86.9583937 C58.1842268,83.7344152 53.8029229,70.715562 60.7532354,53.0843636 C60.7532354,53.0843636 71.5019501,49.6441813 95.9626412,66.2049595 C106.172967,63.368876 117.123047,61.9465949 128.00106,61.8978432 C138.879073,61.9465949 149.837632,63.368876 160.067033,66.2049595 C184.49805,49.6441813 195.231926,53.0843636 195.231926,53.0843636 C202.199197,70.715562 197.815773,83.7344152 196.486756,86.9583937 C204.694018,95.9223682 209.660343,107.340889 209.660343,121.315702 C209.660343,170.478725 179.716133,181.303747 151.213281,184.472614 C155.80443,188.444828 159.895342,196.234518 159.895342,208.176593 C159.895342,225.303317 159.746968,239.087361 159.746968,243.303333 C159.746968,246.709601 162.05102,250.70089 168.53925,249.443941 C219.370432,232.499507 256,184.536204 256,128.00106 C256,57.3066942 198.691187,0 128.00106,0 Z M47.9405593,182.340212 C47.6586465,182.976105 46.6581745,183.166873 45.7467277,182.730227 C44.8183235,182.312656 44.2968914,181.445722 44.5978808,180.80771 C44.8734344,180.152739 45.876026,179.97045 46.8023103,180.409216 C47.7328342,180.826786 48.2627451,181.702199 47.9405593,182.340212 Z M54.2367892,187.958254 C53.6263318,188.524199 52.4329723,188.261363 51.6232682,187.366874 C50.7860088,186.474504 50.6291553,185.281144 51.2480912,184.70672 C51.8776254,184.140775 53.0349512,184.405731 53.8743302,185.298101 C54.7115892,186.201069 54.8748019,187.38595 54.2367892,187.958254 Z M58.5562413,195.146347 C57.7719732,195.691096 56.4895886,195.180261 55.6968417,194.042013 C54.9125733,192.903764 54.9125733,191.538713 55.713799,190.991845 C56.5086651,190.444977 57.7719732,190.936735 58.5753181,192.066505 C59.3574669,193.22383 59.3574669,194.58888 58.5562413,195.146347 Z M65.8613592,203.471174 C65.1597571,204.244846 63.6654083,204.03712 62.5716717,202.981538 C61.4524999,201.94927 61.1409122,200.484596 61.8446341,199.710926 C62.5547146,198.935137 64.0575422,199.15346 65.1597571,200.200564 C66.2704506,201.230712 66.6095936,202.705984 65.8613592,203.471174 Z M75.3025151,206.281542 C74.9930474,207.284134 73.553809,207.739857 72.1039724,207.313809 C70.6562556,206.875043 69.7087748,205.700761 70.0012857,204.687571 C70.302275,203.678621 71.7478721,203.20382 73.2083069,203.659543 C74.6539041,204.09619 75.6035048,205.261994 75.3025151,206.281542 Z M86.046947,207.473627 C86.0829806,208.529209 84.8535871,209.404622 83.3316829,209.4237 C81.8013,209.457614 80.563428,208.603398 80.5464708,207.564772 C80.5464708,206.498591 81.7483088,205.631657 83.2786917,205.606221 C84.8005962,205.576546 86.046947,206.424403 86.046947,207.473627 Z M96.6021471,207.069023 C96.7844366,208.099171 95.7267341,209.156872 94.215428,209.438785 C92.7295577,209.710099 91.3539086,209.074206 91.1652603,208.052538 C90.9808515,206.996955 92.0576306,205.939253 93.5413813,205.66582 C95.054807,205.402984 96.4092596,206.021919 96.6021471,207.069023 Z" fill="#161614"></path></g></svg></div><div class="notion-external-description"><div class="notion-external-title">open-r1</div><div class="notion-external-block-desc">huggingface<span> • </span>Updated Jun 10, 2025</div></div></a><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-20ed0968b24580a1938cf2ffea943817" data-id="20ed0968b24580a1938cf2ffea943817"><span><div id="20ed0968b24580a1938cf2ffea943817" class="notion-header-anchor"></div><a class="notion-hash-link" href="#20ed0968b24580a1938cf2ffea943817" title="有redundant reasoning steps"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">有redundant reasoning steps</span></span></h3><div class="notion-text notion-block-20ed0968b245805ebb50e12059369780">generated long CoT traces often contain many redundant reasoning steps even to the very simple question (Chen et al., 2025; Aggarwal and Welleck, 2025; Yang et al., 2025; Zhang et al., 2025)</div><div class="notion-row notion-block-20ed0968b24580c0970ad98c3ea976d8"><div class="notion-column notion-block-20ed0968b2458070a507e002bf2a3b13" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b245806094fcf9724caeafd9" href="https://arxiv.org/abs/2503.04697"><div><div class="notion-bookmark-title">L1: Controlling How Long A Reasoning Model Thinks With...</div><div class="notion-bookmark-description">Reasoning language models have shown an uncanny ability to improve performance at test-time by ``thinking longer&#x27;&#x27;-that is, by generating longer chain-of-thought sequences and hence using more...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-8060-94fc-f9724caeafd9&amp;t=20ed0968-b245-8060-94fc-f9724caeafd9" alt="L1: Controlling How Long A Reasoning Model Thinks With..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2503.04697</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-8060-94fc-f9724caeafd9&amp;t=20ed0968-b245-8060-94fc-f9724caeafd9" alt="L1: Controlling How Long A Reasoning Model Thinks With..." loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div><div class="notion-column notion-block-20ed0968b245806f9d78cf4dab92e643" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b245803aaf21fdb8262fb661" href="https://arxiv.org/abs/2412.21187"><div><div class="notion-bookmark-title">Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs</div><div class="notion-bookmark-description">The remarkable performance of models like the OpenAI o1 can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-803a-af21-fdb8262fb661&amp;t=20ed0968-b245-803a-af21-fdb8262fb661" alt="Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2412.21187</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-803a-af21-fdb8262fb661&amp;t=20ed0968-b245-803a-af21-fdb8262fb661" alt="Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs" loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div></div><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b2458059be86e5298adbfe13" href="https://arxiv.org/abs/2501.12948"><div><div class="notion-bookmark-title">DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via...</div><div class="notion-bookmark-description">We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-8059-be86-e5298adbfe13&amp;t=20ed0968-b245-8059-be86-e5298adbfe13" alt="DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2501.12948</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-8059-be86-e5298adbfe13&amp;t=20ed0968-b245-8059-be86-e5298adbfe13" alt="DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via..." loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-20ed0968b245806aa006dd194f21bddd">Those redundant reasoning steps may not only bring unnecessary computation burden during test time, but also affect the reasoning performance (Sui et al., 2025; Aggarwal and Welleck, 2025; Wu et al., 2025; Marjanovi ́c et al., 2025)</div><div class="notion-text notion-block-20ed0968b24580c7a51af4189b216418">这些冗余的推理步骤不仅会带来不必要的计算消耗，还会影响推理表现，并且会影响蒸馏过程。</div><div class="notion-row notion-block-20ed0968b2458074af4ec63f1ff3de85"><div class="notion-column notion-block-20ed0968b245804cb9b1c72b6bb06d67" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b24580d3ba02d039b84f67e4" href="https://arxiv.org/abs/2503.16419"><div><div class="notion-bookmark-title">Stop Overthinking: A Survey on Efficient Reasoning for Large...</div><div class="notion-bookmark-description">Large Language Models (LLMs) have demonstrated remarkable capabilities in complex tasks. Recent advancements in Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, have further...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-80d3-ba02-d039b84f67e4&amp;t=20ed0968-b245-80d3-ba02-d039b84f67e4" alt="Stop Overthinking: A Survey on Efficient Reasoning for Large..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2503.16419</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-80d3-ba02-d039b84f67e4&amp;t=20ed0968-b245-80d3-ba02-d039b84f67e4" alt="Stop Overthinking: A Survey on Efficient Reasoning for Large..." loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div><div class="notion-column notion-block-20ed0968b2458060aae8e5f61638dc17" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b24580d494cfe1615071fbe9" href="https://arxiv.org/abs/2503.04697"><div><div class="notion-bookmark-title">L1: Controlling How Long A Reasoning Model Thinks With...</div><div class="notion-bookmark-description">Reasoning language models have shown an uncanny ability to improve performance at test-time by ``thinking longer&#x27;&#x27;-that is, by generating longer chain-of-thought sequences and hence using more...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-80d4-94cf-e1615071fbe9&amp;t=20ed0968-b245-80d4-94cf-e1615071fbe9" alt="L1: Controlling How Long A Reasoning Model Thinks With..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2503.04697</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-80d4-94cf-e1615071fbe9&amp;t=20ed0968-b245-80d4-94cf-e1615071fbe9" alt="L1: Controlling How Long A Reasoning Model Thinks With..." loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div></div><div class="notion-row notion-block-20ed0968b24580d0815cc0fe72e69417"><div class="notion-column notion-block-20ed0968b24580ce9cd4c22f4d51fa92" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b24580fab28ec1e6b968dc8b" href="https://arxiv.org/abs/2503.24370"><div><div class="notion-bookmark-title">Effectively Controlling Reasoning Models through Thinking Intervention</div><div class="notion-bookmark-description">Reasoning-enhanced large language models (LLMs) explicitly generate intermediate reasoning steps prior to generating final answers, helping the model excel in complex problem-solving. In this...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-80fa-b28e-c1e6b968dc8b&amp;t=20ed0968-b245-80fa-b28e-c1e6b968dc8b" alt="Effectively Controlling Reasoning Models through Thinking Intervention" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2503.24370</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-80fa-b28e-c1e6b968dc8b&amp;t=20ed0968-b245-80fa-b28e-c1e6b968dc8b" alt="Effectively Controlling Reasoning Models through Thinking Intervention" loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div><div class="notion-column notion-block-20ed0968b245806b873edfcab37c51cc" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.5)"><div class="notion-row"><a target="_blank" rel="noopener noreferrer" class="notion-bookmark notion-block-20ed0968b245809783a7faa4d193f04d" href="https://arxiv.org/abs/2504.07128"><div><div class="notion-bookmark-title">DeepSeek-R1 Thoughtology: Let&#x27;s think about LLM Reasoning</div><div class="notion-bookmark-description">Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=20ed0968-b245-8097-83a7-faa4d193f04d&amp;t=20ed0968-b245-8097-83a7-faa4d193f04d" alt="DeepSeek-R1 Thoughtology: Let&#x27;s think about LLM Reasoning" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2504.07128</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=20ed0968-b245-8097-83a7-faa4d193f04d&amp;t=20ed0968-b245-8097-83a7-faa4d193f04d" alt="DeepSeek-R1 Thoughtology: Let&#x27;s think about LLM Reasoning" loading="lazy" decoding="async"/></div></a></div></div><div class="notion-spacer"></div></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-20ed0968b24580798d5af816b61e662a" data-id="20ed0968b24580798d5af816b61e662a"><span><div id="20ed0968b24580798d5af816b61e662a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#20ed0968b24580798d5af816b61e662a" title="如何解决这个issue"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">如何解决这个issue</span></span></h3><div class="notion-text notion-block-20ed0968b24580e38896ff031e643261">一般是启发式的方法</div><div class="notion-text notion-block-20ed0968b245801d9182fccb0c1b8f5d">minimum reasoning length with correct final answer (Chen et al., 2025), design length based rewards for reinforcement learning (Aggarwal and Welleck, 2025; Yi and Wang, 2025; Yang et al., 2025), or advanced prompting methods (Wu et al., 2025; Munkhbat et al., 2025; Xia et al., 2025; Han et al., 2025; Nayab et al., 2025).</div><div class="notion-text notion-block-20ed0968b2458068b14bd25a9ed5074a">要么依赖于reward的重新设计，或者不考虑目标SLM在选择长COT训练数据时的推理能力。</div><div class="notion-callout notion-gray_background_co notion-block-20fd0968b2458039a800e7822846f1d4"><div class="notion-page-icon-inline notion-page-icon-span"><span class="notion-page-icon" role="img" aria-label="💡">💡</span></div><div class="notion-callout-text"><div class="notion-text notion-block-20fd0968b24580aaa9afd138150bf53d"><b>How can high-quality CoT traces generated by large reasoning models be efficiently distilled into SLMs?</b></div></div></div><div class="notion-blank notion-block-20fd0968b2458008b825f74f28115d39"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[ACL2025 SLM]]></title>
            <link>http://preview.tangly1024.com/article/ACL2025SLM</link>
            <guid>http://preview.tangly1024.com/article/ACL2025SLM</guid>
            <pubDate>Tue, 10 Jun 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-20ed0968b245805e9ba3dc58120cc011"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><div class="notion-text notion-block-20ed0968b2458062b321f78260544b82"><b>
ACL 2025小型语言模型推理研究进展分析</b>

<b>
I. 引言：ACL 2025小型语言模型推理研究的演进格局</b>

<b>
A. 小型语言模型的兴起与推理的必要性</b>
近年来，人工智能领域见证了向小型语言模型（SLM）发展的显著趋势。这一转变的背后有多重驱动因素，包括对更高效率、更广泛可及性、更低计算成本以及边缘设备部署适用性的追求 <b>1</b>。然而，仅仅缩小模型尺寸不足以释放SLM的全部潜力。推理能力，作为超越简单模式匹配或文本生成的关键，使SLM能够执行复杂任务、深入理解上下文并进行更智能的交互，因此成为SLM研究的核心议题。作为自然语言处理领域的顶级会议，ACL 2025汇聚了该领域的前沿研究，为我们展示了SLM推理方面最新的进展和贡献。
<b>
B. ACL 2025 SLM推理研究主题概览</b>

纵观ACL 2025关于SLM推理的相关论文，可以观察到几个主要的研究方向。这些方向共同描绘了当前SLM推理研究的全貌：
• <b>基础理解与基准测试：</b> 对SLM推理能力进行系统性评估，明确其当前水平、优势与局限。
• <b>架构创新与压缩技术：</b> 研发专为推理任务优化的SLM架构和高效压缩方法。
• <b>知识迁移与蒸馏：</b> 从大型语言模型（LLM）向SLM迁移知识，以赋予SLM更强的推理能力。
• <b>新型推理框架：</b> 提出创新的框架（如智能体框架、模块化框架、协作式框架）以增强或实现SLM的复杂推理。
• <b>数据的作用：</b> 探究训练数据，特别是合成数据，在培养SLM推理能力中的角色。
• 特定领域应用： 在具体应用场景中展示SLM的推理能力，验证其实用价值。
这些研究方向表明，学术界正从多个维度探索提升SLM推理能力的途径，并非依赖单一解决方案，而是呈现出多技术融合的趋势。
<b>
C. 报告范围与结构</b>

本报告旨在综合分析ACL 2025主要会议中与SLM推理相关的若干研究论文。报告将首先概述这些论文的核心贡献，随后分章节深入探讨SLM推理的基础理解、架构创新、知识迁移、新型框架以及实际应用等关键方面。最后，报告将总结当前研究的主要趋势，并展望未来可能的研究方向。
<b>
D. ACL 2025 SLM推理相关论文概览表</b>

为了清晰展示本报告所分析的论文，下表总结了这些论文的标题、主要作者（若可从资料中可靠获取）、其在SLM推理领域的主要关注点、关键技术或贡献，以及相关的资料来源编号。论文标题 (Paper Title)主要作者 (Lead Author(s))SLM推理相关主要关注点 (Primary Focus related to SLM Reasoning)关键技术/贡献 (Key Techniques/Contributions)相关资料ID (Relevant Snippet ID(s))HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval未明确通过联合训练实现多模态检索中的SLM推理联合多任务训练，多模态推理<b>2</b>TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs未明确针对领域特定推理的高效SLM架构渐进式层丢弃，基于激活值的度量<b>2</b>Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models未明确 (IIPL CAU Lab)通过LLM集成增强SLM推理插件模块，微调策略<b>2</b>Demystifying Small Language Models for Edge Deployment未明确SLM推理基准测试，识别局限性与优化路径SLM综合研究，上下文学习能力评估<b>2</b>A Strategic Coordination Framework of Small LMs Matches Large LMs in Data SynthesisXin Gao 等协作式SLM推理生成高质量数据（推理任务的基础）多智能体 (生成器、评审器、裁决器) SLM框架<b>2</b>DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based DistillationJennifer Chen 等将RAG能力蒸馏到SLM以实现基于事实和证据的推理基于证据和知识图谱的蒸馏<b>2</b>Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu 等先进的LLM推理框架，对SLM开发具有启示意义工具使用，思维导图智能体，网络搜索，代码执行<b>2</b>PCoT: Persuasion-Augmented Chain of Thought for Detecting Fake News and Social Media Disinformation未明确在SLM中应用CoT式推理进行虚假新闻检测说服增强的思维链<b>2</b>LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMsChansung Park 等将LLM推理能力迁移到SLM以进行本地部署LLMOps，合成数据生成，迭代式微调<b>2</b>A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks AlignmentJean-Philippe Corbeil 等通过模块化和合成数据开发领域特定（临床）SLM推理能力预指令微调，模型合并，合成指令数据集 (MediFlow)<b>2</b>Flipping Knowledge Distillation: Leveraging Small Models’ Expertise to Enhance LLMs in Text Matching未明确SLM增强LLM，可能实现双向推理能力提升知识蒸馏 (SLM到LLM)，候选标注<b>2</b>ACL这样的顶级会议上出现大量关注SLM的研究，标志着一个重要范式的成熟。SLM不再仅仅是LLM的“轻量版”，而已发展成为一个独特且关键的研究领域，拥有其自身特有的挑战和创新方向，尤其是在复杂的推理领域。这一现象的出现，源于用户查询明确指向ACL 2025，且会议论文列表中包含了众多在标题或摘要中提及“小型语言模型”或“SLM”的论文，其中许多与推理或促成推理相关 <b>2</b>。历史上，语言模型研究的突破多集中于规模扩展（LLM）。如今，SLM研究，特别是关于推理等高级能力的研究，在顶级会议上的显著地位表明，学术界已充分认识到SLM的重要性及其面临的独特挑战。这不仅仅关乎缩小模型体积，更关乎如何使小型模型更“智能”，具备复杂的认知能力。对SLM推理能力的追求，在很大程度上受到实际部署需求的驱动，例如边缘计算、隐私保护和成本控制，这一点在诸如《Demystifying Small Language Models for Edge Deployment》<b>1</b> 和《LlamaDuo》<b>12</b> 等论文中得到了体现。这些研究表明，SLM推理的研究并非纯粹的学术探索，而是与现实世界的应用可行性及人工智能的普及化紧密相连。例如，《Demystifying SLMs》明确提及“资源受限设备”、“边缘部署”以及对“高效普适部署”的需求 <b>1</b>。而《LlamaDuo》则通过将能力迁移到本地SLM来解决操作依赖、隐私顾虑和离线需求等问题 <b>12</b>。推理是使LLM如此强大的核心能力之一。若要SLM在这些受限环境中真正发挥作用，它们必须具备一定程度的推理能力。因此，实际应用中的制约因素直接推动了提升SLM推理能力的研究议程，其更广泛的影响是推动人工智能向更易获取、更注重隐私的方向发展。
<b>
II. 基础理解：SLM推理的基准测试与深入剖析</b>

<b>
A. 建立基线：《Demystifying Small Language Models for Edge Deployment》研究</b>
在SLM推理研究领域，确立坚实的基线和全面的理解至关重要。《Demystifying Small Language Models for Edge Deployment》一文 <b>2</b> 在这方面做出了重要贡献。该研究对超过60个公开可用的SLM（如Microsoft Phi和Google Gemma）进行了首次全面的调查。其与推理的相关性不言而喻，因为它评估了模型的通用任务性能，其中就包括了常识推理和上下文学习（in-context learning, ICL）能力。一个核心发现是，当前最先进的SLM在通用任务上的表现甚至可以超越参数量达到7B（70亿）的模型，这证明了它们在实际应用中的可行性 <b>1</b>。这一数据点意义重大，它表明参数数量并非决定推理能力的唯一因素，尤其对于特定类型的推理任务而言。然而，该研究也明确指出了SLM在上下文学习能力方面存在的局限性 <b>1</b>。ICL通常被视为更复杂推理能力的一个组成部分或代理指标，因此这一局限性对于SLM推理而言是一个关键瓶颈。针对推理能力的提升，该论文确定了一些关键的优化方向，例如动态任务特定路由（可能将任务路由到SLM内部专门的推理模块）和架构-硬件协同设计 <b>1</b>。词汇表/KV缓存压缩虽然主要目标是提升效率，但通过在相同资源预算下支持更长的上下文或容纳更多参数，也间接支持了更复杂的推理过程。
这项工作提供了一个基础性的全局视角，确立了基准，并指出了SLM在推理方面的具体弱点（如ICL）和潜在优势，从而为未来SLM推理研究指明了方向。它为其他专门的推理增强技术提供了可以构建和评估的“事实基础”。
<b>
B. 其他基础层面问题</b>
从SLM的总体发展趋势推断，还存在其他一些基础性问题。例如，需要深入探讨SLM目前能够处理的推理类型（例如，简单的演绎推理、溯因推理，相较于复杂的多跳推理或抽象推理）。此外，如何有效评估SLM的推理能力也是一个挑战，特别是当许多基准测试最初是为LLM设计的时候。尽管并非专门针对推理，但《Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation》一文 <b>2</b> 也指出了对定制化评估方法的需求。尽管SLM展现出了一定的潜力（例如，“超越7B模型”），但其“有限的上下文学习能力”<b>1</b> 对实现复杂的推理构成了显著瓶颈。上下文学习通常被认为是模型涌现推理能力的途径之一。这一观察结果表明，SLM推理能力的突破可能需要采用与LLM的ICL根本不同的方法，或者开发针对SLM架构高度优化的ICL技术。具体而言，《Demystifying SLMs》论文 <b>1</b> 同时强调了SLM在通用任务上的惊人能力以及它们在ICL方面的特定弱点。在LLM中，强大的ICL能力通常与更复杂的推理能力（例如，对新任务的少样本推理）相关联，甚至是其先决条件。如果SLM在ICL方面表现不佳，它们可能难以从少量示例中归纳推理模式，而这正是高效学习和推理的一个标志。这意味着简单地缩小LLM的架构和训练方法可能不足以在SLM中实现稳健的推理；需要针对性地解决SLM ICL局限性或绕过这些局限性的新颖技术。论文中提出的“动态任务特定路由”和“架构-硬件协同设计”的建议 <b>1</b>，暗示了未来SLM的推理能力可能并非来自单一的、整体式的SLM，而是源于由专门化的SLM或SLM组件构成的、针对特定硬件优化的协同系统。这一思路与《A Strategic Coordination Framework》以及模块化方法中的主题相呼应。《Demystifying SLMs》<b>1</b> 将“动态任务特定路由”作为一个优化方向。推理本身并非单一任务，而是多种子任务的集合（例如，逻辑演绎、因果推断、规划）。一个小型、整体式的模型可能难以同时擅长所有这些推理子类型。动态路由可以允许SLM系统激活针对特定推理类型优化的不同专用路径，甚至不同的微模型，所有这些都在一个较小的资源占用内完成。这与<b>2</b>中许多LLM论文中出现的“专家混合”（MoE）思想（尽管并非所有都直接针对SLM，但原理是相关的）以及《A Strategic Coordination Framework》<b>6</b> 中多个小型LM协作的思想相关联。这意味着为了实现高级推理，SLM系统正朝着更复杂、异构的方向发展。
<b>
III. SLM推理的架构创新与效率提升</b>

<b>
A. 为思考而压缩：“TrimLLM”方法</b>
《TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs》一文 <b>2</b> 直接关注SLM的效率问题，这对于部署推理能力至关重要。该方法的核心思想是在微调过程中采用渐进式的层丢弃策略。它利用校准数据集和基于激活值的度量标准来识别并移除非必要的网络层 <b>3</b>。这一策略基于“层级特化”（layer-wise specialization）的假设，即不同的层对特定领域的知识贡献度不同 <b>3</b>。通过创建更紧凑的模型（可能小于原始尺寸的50%），同时在特定领域保持相近的性能 <b>17</b>，TrimLLM使得在有限的计算预算内执行更复杂的推理过程成为可能。它实现了“无论硬件和深度学习框架如何，都能加速推理” <b>18</b>。
这对于设备端推理至关重要，因为在这些场景中，内存和速度是首要考虑的因素。该研究表明，SLM在特定领域的推理能力可以得到高度优化。<b>
B. 2</b>中其他以效率为中心的技术（可能影响推理）虽然<b>2</b>中许多关于专家混合（MoE）的论文主要针对LLM，但稀疏激活和专门化专家的原则（例如，《Accelerating Dense LLMs via L0-regularized Mixture-of-Experts》、《DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts》、《STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning》）可以被调整或正在被探索用于SLM，以在不完全牺牲推理深度的情况下提高效率。诸如《Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models》和《MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts》<b>2</b> 等量化方法也有助于将能力更强（因此可能具有更好推理能力）的模型压缩到更小的体积中。《P$^2$ Law: Scaling Law for Post-Training After Model Pruning》<b>2</b> 的研究可能为模型剪枝后如何保持推理能力提供指导。诸如TrimLLM <b>3</b> 这样的技术突显了模型架构、领域特异性和推理能力之间的关键相互作用。“层级特化”的发现表明，为了使SLM在特定领域有效推理，大型预训练模型中的并非所有部分都同等重要。这使得有针对性的压缩成为可能，从而保留甚至磨练特定应用的推理能力。TrimLLM的成功依赖于“层级特化” <b>3</b>——即某些层对于特定领域更为关键。推理通常是领域相关的（例如，医学推理与金融推理）。通过识别并移除与目标领域推理模式关联较小的层，TrimLLM可以创建一个更小、更专业化且在该领域推理任务上更高效的模型。这揭示了一个因果联系：理解各层对领域知识的贡献度，可以实现高效压缩，进而使得在SLM上进行实用的领域特定推理成为可能。TrimLLM专注于移除整个网络层，而其他方法如MoE则致力于实现层内的稀疏性或通过仅激活部分专家单元来达到目的。这两者之间存在潜在的协同作用（例如，将层丢弃与在剩余层中采用MoE结构相结合），同时也构成了一个比较点：对于SLM而言，哪种方法或方法的组合能够在每个参数/每秒浮点运算次数（FLOPs）下产生最佳的推理能力？TrimLLM本身指出，它与其他模型压缩技术是“正交的”，并且可以组合使用以达到更高的压缩率 <b>18</b>。这并非直接的矛盾，而是指向了实现效率的不同理念。更广泛的意义在于，需要进行比较研究，探讨这些不同的压缩/效率策略具体如何影响SLM中各种
<em>类型</em>的推理。架构创新（如TrimLLM、量化、剪枝以及潜在的SLM-MoE）带来的效率提升，是更复杂的推理算法（如后续将讨论的RAG或智能体方法）能够在SLM上被考虑的关键<em>促成因素</em>。没有这种基础效率，小型设备上的高级推理将无从谈起。高级推理通常意味着更多的计算步骤或需要访问更多的知识。而SLM根据其定义，是资源受限的。像TrimLLM这样的技术减少了模型的深度和计算成本 <b>18</b>。这种减少在受限的资源预算内创造了“余量”。然后，这个余量可以用于更复杂的推理算法（例如，RAG中检索步骤的开销，或智能体框架中工具调用的开销），否则这些算法对于未经优化的SLM来说成本过高。因此，架构效率是SLM上部署高级推理范式的直接促成因素。
<b>
IV. 弥合差距：SLM推理的知识迁移与增强</b>

<b>
A. 从LLM到SLM的推理能力蒸馏</b>

将LLM中蕴含的复杂推理模式迁移到SLM是提升后者能力的关键途径。ACL 2025的多项研究聚焦于此。<b>《DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation》</b> <b>2</b> 是该领域的核心工作之一。该研究针对LLM驱动的检索增强生成（RAG）系统虽然功能强大但资源消耗大且易产生幻觉的问题，以及SLM虽小巧但需要此类高级能力的需求，提出了DRAG框架。DRAG的核心思想是将大型教师RAG模型的RAG能力蒸馏到小型的学生SLM中。它采用基于证据和知识图谱的蒸馏方法，以确保事实准确性并减轻幻觉 <b>8</b>。具体而言，教师模型首先生成与输入问题相关的证据和知识图谱，然后学生SLM学习模仿这种由证据驱动的推理过程 <b>8</b>。DRAG旨在直接赋予SLM进行检索增强生成的能力，这是一种涉及访问和综合外部知识的复杂推理任务，从而使SLM的推理更加基于事实且可靠。实验表明，DRAG的性能比先前的MiniRAG等方法提升高达27.7% <b>8</b>。<b>《LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs》</b> <b>2</b> 则侧重于将包括推理在内的知识和能力从大型服务型LLM平滑迁移到本地SLM的实践流程。该方法通过使用由大型服务LLM生成的合成数据集对SLM进行微调。如果微调后的SLM性能未达预期，则会利用服务LLM生成的额外相似数据进行进一步的迭代微调，以确保SLM在特定下游任务上的能力能够达到甚至超越服务LLM的水平 <b>12</b>。这项工作为创建具有针对性推理能力的SLM提供了一条途径，以满足特定下游任务的需求，同时保障服务连续性并解决隐私和离线运行等问题。这意味着SLM可以通过训练来复制大型模型在专门应用中的推理输出。此外，<b>《Towards the Law of Capacity Gap in Distilling Language Models》</b> <b>2</b> 虽然摘要信息有限，但其标题表明该研究探讨了知识蒸馏中能力差距的规律，这对于理解和优化推理能力的迁移至关重要。
<b>
B. 协同方法：插件、微调与双向知识流</b>

除了直接蒸馏，研究者们也在探索通过插件、微调以及更复杂的知识交互模式来增强SLM的推理能力。<b>《Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models》</b> <b>2</b> 由IIPL CAU实验室提出 <b>4</b>，旨在弥合SLM与LLM之间的能力差距，这其中必然包含推理能力。尽管具体方法在现有资料中未详述，但其标题暗示了通过使用插件模块（可能是专门的推理单元）和有针对性的微调策略（可能由LLM指导）来增强SLM。<b>《Flipping Knowledge Distillation: Leveraging Small Models’ Expertise to Enhance LLMs in Text Matching》</b> <b>2</b> 提出了一个有趣的视角，对传统的LLM到SLM的蒸馏方向进行了反转或补充。根据相关资料描述的CanDist框架 <b>14</b>，其核心思想是引导LLM提供
<em>候选标注</em>（即多个可能的标签），而非单一的黄金标准标签。然后，SLM从这些候选标注中进行蒸馏，并使用一个分布精炼机制。这表明SLM可以帮助提炼或更好地利用LLM生成的知识/数据。尽管该研究主要聚焦于文本匹配，但文本匹配任务本身也常涉及细致的推理过程。如果SLM能够从LLM的输出中提炼出更稳健的信号，或者SLM本身在效率或特定领域拥有独特的“专长”可以反过来优化LLM的流程，那么这将指向一个更复杂的生态系统，其中推理能力是共同发展的。知识蒸馏是赋予SLM达到LLM级别推理能力的主要手段。然而，这一技术正从简单的输出复制，演变为更复杂的方法，例如DRAG中基于证据和知识图谱的蒸馏 <b>8</b>，或LlamaDuo中基于合成数据的迭代优化 <b>12</b>。这表明，对于复杂的推理任务而言，朴素的蒸馏方式已不足够。多项关键研究（如DRAG、LlamaDuo）明确采用蒸馏或LLM生成的据来训练SLM。DRAG <b>8</b> 强调的不仅仅是模仿输出词元，更是RAG的整个<em>过程</em>，即基于证据和知识图谱的蒸馏。LlamaDuo <b>12</b> 则在SLM性能不足时，采用“迭代过程”和“额外的相似数据”进行改进，这意味着一种更具指导性和精细化的蒸馏。这反映了一个发展过程：早期的蒸馏可能侧重于任务性能，但对于推理而言，得出答案的
<em>方法</em>本身也需要被蒸馏，这便要求采用更结构化的途径。一个新兴的主题是知识流动的双向性或协同性。《Flipping Knowledge Distillation》<b>2</b> 的出现，暗示了知识流动可能从纯粹的单向（LLM -&gt; SLM）模式转变。SLM或许能在提炼LLM输出或流程方面发挥作用，从而形成一种共生关系。这种关系可能间接提升最终被蒸馏<em>回</em>SLM的知识质量，或用于创建更优的SLM。例如，CanDist框架 <b>14</b> 讨论了SLM从LLM提供的“候选标注”中进行蒸馏，从而提炼学习信号。如果SLM能帮助LLM产生更好（例如，更细致或更可靠）的输出或数据，那么这个经过改进的LLM就能成为其他SLM的更优质教师。这可能形成一个正反馈循环，或一个更细致的生态系统，其中SLM不再仅仅是被动的接受者，而是知识库的积极贡献者，这些知识随后可用于更有效的SLM推理。LlamaDuo的LLMOps流水线的成功 <b>12</b>，突显了稳健的工程实践和MLOps对于从LLM有效创建具备推理能力的SLM的重要性。这不仅仅关乎算法本身，更关乎数据生成、训练、评估和迭代的整个流程。LlamaDuo明确是一个“LLMOps流水线” <b>12</b>，它包含由LLM生成合成数据、对SLM进行微调、性能评估以及迭代改进等环节。这些都是成熟MLOps周期的组成部分。要成功地将像推理这样的复杂能力从大型通用模型迁移到小型专用模型，需要对这些步骤进行审慎管理。这意味着随着SLM推理技术变得日益复杂，支持其开发和部署的工程基础设施也将愈发关键。
<b>
V. SLM复杂推理的新型框架与方法论</b>

<b>
A. 智能体推理及其在SLM中的潜力</b>
<b>《Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools》</b> <b>2</b> 虽然主要在LLM（如DeepSeek-R1 <b>26</b>）上进行了演示，但其框架对SLM具有重要的启示意义。该框架的核心思想是通过集成外部工具使用型智能体（如网络搜索、代码执行）和一个结构化的“思维导图”（Mind Map）记忆模块（一个用于追踪逻辑关系的知识图谱）来增强LLM的推理能力 <b>10</b>。这使得模型能够在复杂的多步骤推理过程中动态检索信息并管理结构化上下文。对于SLM而言，挑战在于如何将这些智能体能力蒸馏或适配到其较小的模型结构中。如果成功，SLM将能够通过学习<em>何时</em>以及<em>如何</em>调用外部工具或访问结构化记忆，来克服其固有的知识局限性，从而执行超越其内部参数所能支持的推理。这与《Demystifying SLMs》中提出的“动态任务特定路由”思想不谋而合。该论文还提及，可以利用工具调用频率作为启发式方法进行“测试时扩展”（test-time scaling）<b>10</b>，这可能是一种计算成本较低的提升SLM推理输出质量的途径。
<b>
B. 协作式推理：多个SLM的战略协调</b>
<b>《A Strategic Coordination Framework of Small LMs Matches Large LMs in Data Synthesis》(GRA)</b> <b>2</b> 提出了一种新颖的方式，使多个SLM能够协同工作，达到与LLM相媲美的效果，特别是在数据合成方面——而高质量数据是训练具备良好推理能力模型的基础。GRA框架 <b>6</b> 的设计灵感来源于同行评审流程，其中多个小型LM扮演不同角色——生成器（Generator）、评审器（Reviewer）和裁决器（Adjudicator）——通过迭代优化来精炼数据。高质量、多样化且可靠的数据对于训练模型进行有效推理至关重要。如果一个<em>由SLM组成的协调小组</em>能够生成与大型LLM（例如Qwen-2.5-72B-Instruct <b>6</b>）相当甚至更高质量的数据，这将极大地推动以推理为中心的数据集的民主化创建。这也暗示着，像同行评审和裁决这样复杂的推理过程可以被分解并由专门的SLM处理。
<b>
C. 面向领域特定SLM推理的模块化设计</b>
<b>《A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment》</b> <b>2</b> 展示了如何构建高性能、特定领域（临床）的SLM。该方法论包括在特定医学语料库上对专家SLM进行预指令微调，然后进行模型合并（以统一专家模型并恢复基础能力），最后使用一个大规模的合成指令数据集（MediFlow，包含250万条指令，覆盖多种临床任务和文档类型）进行对齐 <b>13</b>。这种模块化方法在医学实体识别、放射学报告分析和ICD-10编码等需要细致临床推理的任务上取得了显著改进，在某些情况下甚至优于GPT-4 <b>13</b>。其核心在于，模块化（先创建专门的专家模型然后合并）和高质量的合成数据是解锁SLM在复杂领域强大推理能力的关键。
<b>
D. 通过增强提示进行推理：SLM中的思维链</b>
<b>《PCoT: Persuasion-Augmented Chain of Thought for Detecting Fake News and Social Media Disinformation》</b> <b>2</b> 明确地将思维链（Chain of Thought, CoT）的一种变体应用于SLM（或至少是以一种可适配于SLM的方式应用于LLM），以解决一个推理密集型任务。该研究的重点是检测虚假新闻和不实信息，这需要分析声明、证据以及潜在的说服性操纵。CoT提示有助于模型阐明逐步的推理过程。“说服增强”（Persuasion-Augmented）则表明对CoT进行了改进，以考虑文本中的说服性元素，为推理过程增加了另一个维度。在SLM中有效应用CoT的挑战在于，与LLM相比，SLM处理长推理链的能力有限。资料<b>30</b>（描述另一篇关于挖掘虚假新闻中隐含观点以进行检测的论文）和<b>31</b>（关于金融领域CoT用于FMD挑战赛）暗示了使用CoT进行复杂推理任务的更广泛趋势，PCoT很可能在此基础上针对SLM进行了构建。智能体 <b>10</b>、协作式（GRA <b>6</b>）以及模块化 <b>13</b> 框架的出现，表明SLM的研究正从将SLM视为单一的、自足的推理者，转向通过外部化知识/工具、与其他SLM协作或由专门模块组成来放大其推理能力。这与试图将所有推理能力塞进一个小型模型中的传统思路有显著不同。智能体推理 <b>10</b> 明确利用外部工具和记忆。GRA框架 <b>6</b> 则采用多个SLM协同工作的方式。临床SLM论文 <b>13</b> 使用了专家模型的合并。传统的SLM是参数有限的单一模型。这些新框架承认了单个SLM固有的能力限制，并提出通过以下方式克服这些限制：a) 允许SLM
<em>访问</em>外部资源（智能体式），b) <em>划分</em>推理任务（GRA、模块化），或 c) 从专门单元<em>组合</em>技能（模块化）。这意味着未来SLM的高级推理很可能涉及“SLM系统”或集成到更大信息生态系统中的SLM。GRA框架 <b>6</b> 和临床SLM论文（通过MediFlow数据集 <b>13</b>）都强调了高质量、特别是合成数据在开发SLM推理能力中的关键作用。GRA展示了SLM能够
<em>创建</em>这类数据，而临床SLM的研究则表明，SLM可以利用这类数据进行有效<em>对齐</em>。推理，尤其是在专业领域或复杂任务中，需要接触到多样化且准确的推理过程示例或任务演示。通过人工标注大规模获取此类数据成本高昂且困难。而合成数据生成（无论是通过LLM，还是像GRA那样由SLM协作完成）提供了一种可扩展的解决方案。这揭示了一个潜在的因果循环：更好的数据（可能是合成的）-&gt; 更好的SLM推理能力 -&gt; （可能）SLM能够生成质量更高的数据（如GRA所示）。诸如《TrimLLM》（针对特定领域的层剪枝）和《A Modular Approach for Clinical SLMs》等研究有力地表明，在SLM中实现强大推理能力，通过专注于狭窄领域的特化可能比追求LLM所期望的广泛的、类人通用推理更为可行。TrimLLM <b>3</b> 强调通过在“相关数据集”上进行微调，并剪除对“目标领域”贡献较小的层，来使LLM适应“特定任务”。临床SLM论文 <b>13</b> 则创建了针对“相关医学和临床语料库”的“专家模型”，并针对“临床任务”进行对齐。SLM的能力有限，试图用这种有限的能力成为一个通用推理者极具挑战性。然而，通过将这种有限的能力集中于特定领域或一小组推理任务，SLM可以实现高性能，在某些专业情况下甚至超越大型模型（例如，临床SLM在ICD-10编码任务上优于GPT-4 <b>13</b>）。这为SLM推理提供了一个战略方向：为不同领域和任务开发高度专业化的“专家SLM”，而非追求一个万能的SLM。这也与动态任务特定路由的思想相联系。
<b>
VI. SLM推理的应用与评估实践</b>

<b>
A. SLM的多模态推理</b>
<b>《HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval》</b> <b>2</b> 探索了小型和大型语言模型的联合训练在多模态任务中的应用。在该研究中，高效的多模态酒店检索任务可能涉及到跨越文本和图像模态的推理，以便理解用户查询、匹配酒店特征并对选项进行排序。“联合多任务训练”可能是将推理技能迁移到或在SLM组件中共同开发这些技能的一种方式。这项工作表明，SLM正被考虑用于处理复杂的现实世界任务，这些任务需要整合来自多个来源的信息并进行推理。
<b>
B. SLM推理在信息完整性保障中的应用</b>
<b>《PCoT: Persuasion-Augmented Chain of Thought for Detecting Fake News and Social Media Disinformation》</b> <b>2</b> 将CoT风格的推理应用于虚假新闻检测这一关键任务，这可能在SLM中实现。这项任务需要细致的理解、推断，并可能需要追踪逻辑一致性或识别操纵性语言——这些都是推理能力的体现。如果PCoT方法对SLM有效，它将展示SLM在打击错误信息方面的实用价值，而这是一项对复杂推理能力要求很高的任务。
<b>
C. 更广泛评估带来的启示</b>
回顾**《Demystifying Small Language Models for Edge Deployment》** <b>1</b> 的发现，SLM在常识推理基准测试上的表现及其在上下文学习方面的局限性，直接影响了它们在处理新情境时的推理能力。尽管并非专门针对SLM，<b>《Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation》</b> <b>2</b> 指出了改进评估方法学的必要性。对于SLM推理而言，这意味着可能需要开发专门针对SLM能力和典型用例的基准测试，而不是简单地使用缩减版的LLM基准。SLM在多模态检索（HotelMatch-LLM <b>2</b>）和虚假新闻检测（PCoT <b>2</b>）等任务中的应用，表明研究正推动SLM超越基础NLP任务，进入需要更复杂、上下文感知推理的领域。例如，HotelMatch-LLM <b>2</b> 处理多模态酒店检索，这需要理解用户需求（文本/语音）、酒店的视觉方面（图像）并进行匹配——这是一个推理过程。PCoT <b>2</b> 则针对虚假新闻检测，需要分析声明、证据和说服策略。这些并非简单的分类或生成任务，它们涉及整合多条信息并进行推断。SLM在这些应用中的探索显示出一种雄心，即赋予它们与先前LLM领域相当或足以胜任的推理技能。尽管SLM正被应用于推理任务，但《Demystifying SLMs》<b>1</b> 指出了其局限性（例如在ICL方面）。再结合对“定制化基准”的普遍呼吁 <b>2</b>，这暗示了一个潜在的差距：当前的基准可能无法完全捕捉SLM推理的细微差别，或者SLM可能在某些特定<em>类型</em>的推理上表现出色，而这些类型并未被以LLM为中心的评估很好地覆盖。《Demystifying SLMs》<b>1</b> 使用了现有基准，但也强调了SLM在ICL等方面的特定弱点。《Beyond One-Size-Fits-All》<b>2</b> 则普遍主张采用定制化基准。如果SLM正朝着专业化方向发展（如TrimLLM或临床SLM论文所示），那么通用推理基准可能无法反映它们在其特定领域的真实能力。反之，这些通用基准也可能揭示出仅靠专业化无法克服的基础推理缺陷。这意味着需要为SLM推理制定更细致的评估策略，可能包括特定领域的基准，或能够有效测试SLM上可高效执行的特定推理原语的基准。《Demystifying SLMs》<b>1</b> 将其研究范围定义为参数量在1亿到50亿之间的模型。这个范围内的某些“SLM”已经相当强大。通过各种研究所展示的SLM推理能力的持续改进表明，对于一个“小型”模型而言，何为“足够好”的推理能力的标准正在不断提高。<b>1</b>中SLM的定义（1亿-50亿参数）覆盖了广泛的范围，一个50亿参数的模型比一个1亿参数的模型能力要强得多。像DRAG <b>8</b> 和LlamaDuo <b>12</b> 这样的研究正在成功地将能力从更大的LLM迁移到这个SLM范围内的模型。随着这些迁移技术的改进，以及像TrimLLM这样的架构创新使SLM更高效，对“SLM”的基线推理能力预期将会提高。这意味着“小型”和“大型”之间的区别可能更多地关乎部署环境（边缘vs云端），而非固定的推理能力差距，尤其对于专业化任务而言。
<b>
VII. 综合分析：SLM推理的关键洞察与未来轨迹</b>

<b>
A. ACL 2025主要主题与重大突破回顾</b>

ACL 2025的研究成果清晰地勾勒出SLM推理领域的发展脉络：
• <b>效率是先决条件：</b> 架构创新（如TrimLLM <b>3</b>）和对SLM性能的基础性理解（如《Demystifying SLMs》<b>1</b>）为实现实用的SLM推理铺平了道路。效率的提升使得在资源受限的SLM上运行复杂的推理算法成为可能。
• <b>巧妙利用LLM：</b> 复杂的蒸馏技术（如DRAG <b>8</b>，LlamaDuo <b>12</b>）和增强方法（如《Plug-in and Fine-tuning》<b>2</b>）超越了简单的模仿，致力于迁移结构化的推理过程，使SLM能够继承LLM的部分推理能力。
• <b>SLM认知新范式：</b> 研究趋势显示，SLM正从单一、孤立的模型向更开放和协作的系统转变。智能体式（Agentic Reasoning <b>10</b>）、协作式（GRA <b>6</b>）和模块化（临床SLM <b>13</b>）等新方法的出现，使SLM能够执行单个小型孤立模型难以完成的复杂推理。
• <b>数据的力量：</b> 高质量数据，特别是合成数据（如临床SLM研究中的MediFlow <b>13</b>，以及GRA框架 <b>6</b>），在训练和对齐SLM以执行推理任务方面发挥着至关重要的作用。
<b>
B. 已识别的挑战与开放性研究问题</b>

尽管取得了显著进展，SLM推理领域仍面临诸多挑战：
• <b>泛化性与特化性的权衡：</b> 如何在SLM中平衡对广泛推理技能的需求与领域/任务特化带来的实际优势？
• <b>蒸馏推理的鲁棒性：</b> SLM通过蒸馏学到的推理能力在面对分布外输入或对抗性攻击时的鲁棒性如何？诸如《PIG: Privacy Jailbreak Attack on LLMs》<b>2</b> 的论文暗示了LLM存在的漏洞，这些漏洞也可能适用于SLM。
• <b>可解释性与可信度：</b> 随着SLM执行更复杂的推理，如何确保其推理过程透明且可信，尤其是在临床、虚假新闻检测等关键应用中？
• <b>高级框架的可扩展性：</b> 智能体或多SLM协调框架能否在真正资源受限的边缘设备上高效实现？
• <b>评估指标：</b> 仍然需要更好、更针对SLM的推理能力基准测试。
• <b>伦理考量：</b> 随着SLM推理能力的增强（例如PCoT中的说服能力），相关的伦理问题是什么？
<b>
C. 未来潜在研究方向与SLM角色的演变</b>

基于当前的进展和挑战，未来SLM推理的研究可能朝以下方向发展：
• <b>混合模型：</b> 将符号推理引擎与SLM更紧密地集成，结合两者的优势。
• <b>SLM推理器的终身学习：</b> 使SLM能够从设备上的新数据中持续适应和改进其推理技能。虽然是会议发现论文，但《Multi-Stage LLM Fine-Tuning with a Continual Learning Setting》<b>29</b> 表明了这一趋势。
• <b>SLM间的协作与学习：</b> 在GRA <b>6</b> 等框架的基础上扩展，以实现更复杂的分布式推理和学习生态系统。
• <b>软硬件协同设计：</b> 进一步研究专为推理任务协同设计SLM架构和硬件加速器 <b>1</b>。
• <b>“智能”边缘：</b> 随着SLM推理能力的提高，边缘设备可能具备更强的自主决策和复杂交互能力，从而改变各个行业。综合来看，ACL 2025的研究成果共同指向一个未来趋势，即复杂推理能力不再局限于大型、中心化的LLM。相反，它正变得日益去中心化，通过功能强大的独立SLM、由工具/协作增强的SLM，或直接嵌入应用和设备中的SLM来实现。边缘部署是一个反复出现的主题 <b>1</b>。LlamaDuo <b>12</b> 明确旨在将能力迁移到本地SLM。智能体推理 <b>10</b> 可以赋予单个SLM访问大量外部知识的能力。GRA框架 <b>6</b> 展示了即便是为推理生成数据也可以在SLM之间实现去中心化。这些集体动向标志着从完全依赖大型云模型到在更靠近数据源或用户的地方实现推理能力的转变，这对隐私、延迟和自主性都具有深远影响。与其说SLM简单地取代LLM，或反之亦然，ACL 2025的研究更揭示了一个日益相互依赖的生态系统。LLM对于引导SLM的推理能力至关重要（通过蒸馏、合成数据生成等方式），但SLM也可能反过来优化LLM的流程（如《Flipping Knowledge Distillation》<b>2</b>），或在由LLM编排的更大工作流中更有效地处理专门任务。许多论文聚焦于LLM到SLM的知识迁移（如DRAG、LlamaDuo）。《Flipping Knowledge Distillation》<b>2</b> 则暗示了SLM对LLM的增强作用。《HotelMatch-LLM》<b>2</b> 采用了小型和大型LM的联合训练。这并非简单的替代关系，而是一种更复杂的相互作用，其中每种类型的模型都利用了对方的优势。LLM提供规模和通用知识；SLM提供效率和专业化。未来很可能涉及混合系统，其中LLM和SLM协同工作，以在不同需求下提供最佳的推理性能。虽然当前的研究主要集中在让SLM<em>执行</em>推理，但对于关键应用而言，一个合乎逻辑的下一步将是使SLM能够<em>对其自身的推理过程进行推理</em>——即具备可解释性、不确定性量化和自我修正的能力。尽管在这些论文中，这尚未成为SLM研究的主导议题，但对可信度的需求暗示了这一发展方向。当前的研究致力于使SLM能够进行推理（例如，PCoT中的CoT，DRAG中的RAG）。随着SLM被部署到临床、虚假新闻检测等敏感领域，仅仅提供答案是不够的；理解SLM<em>如何</em>得出答案以及其<em>置信度如何</em>变得至关重要。这需要元推理能力。虽然LLM研究正在探索这一领域，但对于SLM而言，在能力有限的情况下实现这一点将是一个重大挑战，也是建立信任和确保可靠性的关键研究领域。DRAG框架中对减轻幻觉的关注 <b>8</b> 是朝此方向迈出的早期一步。
<b>
VIII. 结论</b>

ACL 2025展示了小型语言模型（SLM）推理研究领域的蓬勃发展和显著进步。研究工作不再仅仅满足于缩小模型尺寸，而是积极探索如何赋予这些紧凑模型强大的推理能力，以应对日益复杂的现实世界挑战。从对SLM推理能力的<b>基础性理解和基准测试</b>出发，学术界正努力描绘SLM的当前版图，识别其优势与瓶颈，如在通用任务上表现优异但在上下文学习方面仍存局限 <b>1</b>。这为后续研究提供了明确的优化方向。<b>架构创新和效率提升</b>是推动SLM推理实用化的关键。以TrimLLM <b>3</b> 为代表的技术，通过渐进式层丢弃等方法，在保持领域特定性能的同时，显著提升了SLM的推理效率，为在资源受限设备上部署复杂推理应用奠定了基础。<b>知识迁移和增强</b>成为弥合SLM与LLM能力差距的核心策略。DRAG <b>8</b> 等研究通过精密的蒸馏技术，将LLM的检索增强生成等复杂推理能力迁移至SLM，并着力于缓解幻觉问题。同时，LlamaDuo <b>12</b> 等工作则关注于构建稳健的LLMOps流水线，以实现从服务型LLM到本地SLM的平滑能力迁移。更有研究开始探索SLM反哺LLM的可能性 <b>2</b>，预示着一个更加动态和协同的SLM-LLM生态系统。<b>新颖框架和方法论的涌现</b>，如智能体推理 <b>10</b>、多SLM协作（GRA框架 <b>6</b>）以及模块化设计（临床SLM研究 <b>13</b>），正在重新定义SLM执行复杂推理的方式。这些框架通过引入外部工具、分解任务或组合专家模块，使SLM能够突破自身参数限制，处理更深层次的逻辑推演和知识综合。此外，思维链（CoT）等提示工程技术也在被积极探索和应用于SLM，以提升其在特定任务（如虚假信息检测 <b>2</b>）中的推理表现。在<b>应用层面</b>，SLM的推理能力已开始在多模态检索 <b>2</b> 和信息完整性保障等领域得到验证。然而，如何针对SLM的特性设计更有效的
<b>评估基准</b>，仍然是一个亟待解决的问题。
综合来看，ACL 2025的研究揭示了SLM推理的几个核心趋势：对效率和实用性的持续追求；从LLM获取知识和能力的复杂化与精细化；通过创新框架赋予SLM超越个体能力的系统级推理潜能；以及高质量（尤其是合成）数据在驱动SLM推理发展中的核心作用。
未来的研究无疑将继续深化这些方向，同时应对泛化性、鲁棒性、可解释性和伦理等方面的挑战。SLM推理能力的不断突破，预示着一个更加智能、普惠和去中心化的人工智能未来，其中小型模型将在边缘计算、个性化服务和关键行业应用中扮演越来越重要的角色。</div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[FoodSeg]]></title>
            <link>http://preview.tangly1024.com/article/foodseg</link>
            <guid>http://preview.tangly1024.com/article/foodseg</guid>
            <pubDate>Tue, 10 Jun 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[32种中餐食物语义分割模型。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-20dd0968b24580a58229c1e9108518f8"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><a target="_blank" rel="noopener noreferrer" href="https://github.com/smart-diet-assistent/foodseg" class="notion-external notion-external-block notion-row notion-block-20dd0968b24580829effc485e9543ba0"><div class="notion-external-image"><svg viewBox="0 0 260 260"><g><path d="M128.00106,0 C57.3172926,0 0,57.3066942 0,128.00106 C0,184.555281 36.6761997,232.535542 87.534937,249.460899 C93.9320223,250.645779 96.280588,246.684165 96.280588,243.303333 C96.280588,240.251045 96.1618878,230.167899 96.106777,219.472176 C60.4967585,227.215235 52.9826207,204.369712 52.9826207,204.369712 C47.1599584,189.574598 38.770408,185.640538 38.770408,185.640538 C27.1568785,177.696113 39.6458206,177.859325 39.6458206,177.859325 C52.4993419,178.762293 59.267365,191.04987 59.267365,191.04987 C70.6837675,210.618423 89.2115753,204.961093 96.5158685,201.690482 C97.6647155,193.417512 100.981959,187.77078 104.642583,184.574357 C76.211799,181.33766 46.324819,170.362144 46.324819,121.315702 C46.324819,107.340889 51.3250588,95.9223682 59.5132437,86.9583937 C58.1842268,83.7344152 53.8029229,70.715562 60.7532354,53.0843636 C60.7532354,53.0843636 71.5019501,49.6441813 95.9626412,66.2049595 C106.172967,63.368876 117.123047,61.9465949 128.00106,61.8978432 C138.879073,61.9465949 149.837632,63.368876 160.067033,66.2049595 C184.49805,49.6441813 195.231926,53.0843636 195.231926,53.0843636 C202.199197,70.715562 197.815773,83.7344152 196.486756,86.9583937 C204.694018,95.9223682 209.660343,107.340889 209.660343,121.315702 C209.660343,170.478725 179.716133,181.303747 151.213281,184.472614 C155.80443,188.444828 159.895342,196.234518 159.895342,208.176593 C159.895342,225.303317 159.746968,239.087361 159.746968,243.303333 C159.746968,246.709601 162.05102,250.70089 168.53925,249.443941 C219.370432,232.499507 256,184.536204 256,128.00106 C256,57.3066942 198.691187,0 128.00106,0 Z M47.9405593,182.340212 C47.6586465,182.976105 46.6581745,183.166873 45.7467277,182.730227 C44.8183235,182.312656 44.2968914,181.445722 44.5978808,180.80771 C44.8734344,180.152739 45.876026,179.97045 46.8023103,180.409216 C47.7328342,180.826786 48.2627451,181.702199 47.9405593,182.340212 Z M54.2367892,187.958254 C53.6263318,188.524199 52.4329723,188.261363 51.6232682,187.366874 C50.7860088,186.474504 50.6291553,185.281144 51.2480912,184.70672 C51.8776254,184.140775 53.0349512,184.405731 53.8743302,185.298101 C54.7115892,186.201069 54.8748019,187.38595 54.2367892,187.958254 Z M58.5562413,195.146347 C57.7719732,195.691096 56.4895886,195.180261 55.6968417,194.042013 C54.9125733,192.903764 54.9125733,191.538713 55.713799,190.991845 C56.5086651,190.444977 57.7719732,190.936735 58.5753181,192.066505 C59.3574669,193.22383 59.3574669,194.58888 58.5562413,195.146347 Z M65.8613592,203.471174 C65.1597571,204.244846 63.6654083,204.03712 62.5716717,202.981538 C61.4524999,201.94927 61.1409122,200.484596 61.8446341,199.710926 C62.5547146,198.935137 64.0575422,199.15346 65.1597571,200.200564 C66.2704506,201.230712 66.6095936,202.705984 65.8613592,203.471174 Z M75.3025151,206.281542 C74.9930474,207.284134 73.553809,207.739857 72.1039724,207.313809 C70.6562556,206.875043 69.7087748,205.700761 70.0012857,204.687571 C70.302275,203.678621 71.7478721,203.20382 73.2083069,203.659543 C74.6539041,204.09619 75.6035048,205.261994 75.3025151,206.281542 Z M86.046947,207.473627 C86.0829806,208.529209 84.8535871,209.404622 83.3316829,209.4237 C81.8013,209.457614 80.563428,208.603398 80.5464708,207.564772 C80.5464708,206.498591 81.7483088,205.631657 83.2786917,205.606221 C84.8005962,205.576546 86.046947,206.424403 86.046947,207.473627 Z M96.6021471,207.069023 C96.7844366,208.099171 95.7267341,209.156872 94.215428,209.438785 C92.7295577,209.710099 91.3539086,209.074206 91.1652603,208.052538 C90.9808515,206.996955 92.0576306,205.939253 93.5413813,205.66582 C95.054807,205.402984 96.4092596,206.021919 96.6021471,207.069023 Z" fill="#161614"></path></g></svg></div><div class="notion-external-description"><div class="notion-external-title">foodseg</div><div class="notion-external-block-desc">smart-diet-assistent<span> • </span>Updated Jun 11, 2025</div></div></a><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-20dd0968b2458006b870faefc36fdc3b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:712px"><iframe class="notion-asset-object-fit" src="https://yummytanmo.github.io/diet?spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><div class="notion-blank notion-block-20fd0968b24580e08610d8b1b6e227ec"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[A Holistic Lexicon-Based Approach to Opinion Mining]]></title>
            <link>http://preview.tangly1024.com/article/206d0968-b245-8050-b4c8-c7a3395688b5</link>
            <guid>http://preview.tangly1024.com/article/206d0968-b245-8050-b4c8-c7a3395688b5</guid>
            <pubDate>Mon, 02 Jun 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-206d0968b2458050b4c8c7a3395688b5"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-206d0968b24580adb92deb4bba7909fd" data-id="206d0968b24580adb92deb4bba7909fd"><span><div id="206d0968b24580adb92deb4bba7909fd" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580adb92deb4bba7909fd" title="INTRODUCTION"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">INTRODUCTION</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245809fa9c9f1c804c87059" data-id="206d0968b245809fa9c9f1c804c87059"><span><div id="206d0968b245809fa9c9f1c804c87059" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245809fa9c9f1c804c87059" title="背景"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">背景</span></span></h3><div class="notion-text notion-block-206d0968b24580ecbf2cea2d4b3f6c07">商品评价繁杂</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b24580929e80f1853dd3f0ac" data-id="206d0968b24580929e80f1853dd3f0ac"><span><div id="206d0968b24580929e80f1853dd3f0ac" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580929e80f1853dd3f0ac" title="两个任务"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">两个任务</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b245804eb33de76334705c3a"><li>To find product features that have been commented on by reviewers</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b24580138319c2f5122efca2"><li>To decide whether the comments are positive or negative</li></ol><div class="notion-text notion-block-206d0968b245802db87aeba2515c8cc6">本文关注task2 ，we want to oaccurately identify th esemanic orientations of opinions  expressed on each prduct feature by each reviewer. 判断积极、消极、中立。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b2458014b823dce2dae0015f" data-id="206d0968b2458014b823dce2dae0015f"><span><div id="206d0968b2458014b823dce2dae0015f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458014b823dce2dae0015f" title="【13】方法："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">【13】方法：</span></span></h3><div class="notion-text notion-block-206d0968b24580408383cd4c560da980">使用feature附近的opinion words判断对某个特征的态度，通过positive和nagative词语的数量比较简单判断。</div><div class="notion-text notion-block-206d0968b24580a3813eee8d137c5afb"><b>缺点：</b></div><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b24580d285c9ec0c33cdebf8"><li>针对不同语境，词语情感导向不同，专家提供知识库不可取</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b24580cebf2ed16a542de280"></ol><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b2458019a8d3e218f33f6f57" data-id="206d0968b2458019a8d3e218f33f6f57"><span><div id="206d0968b2458019a8d3e218f33f6f57" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458019a8d3e218f33f6f57" title="提出方法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">提出方法</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b24580d69acbc69d2f25092c"><li>运用其他语句、评论的外部信息、证据，推断opinion words的情感导向，不需要前置知识和任何输入</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b2458039bb02f0d9ba37bf10"><li>针对同一句话中相反情感的词语，提出新模型：considering the distance between each opinion word and the product feature. This turns out to be highly effective.</li></ol><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245803fa58fcc1938877c7b" data-id="206d0968b245803fa58fcc1938877c7b"><span><div id="206d0968b245803fa58fcc1938877c7b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245803fa58fcc1938877c7b" title="测试方法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">测试方法</span></span></h3><div class="notion-text notion-block-206d0968b245806683e0fe0b2b4066eb">测试数据</div><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b2458058b030e75f9487b398"><li>bench mark review data set used in [13, 28] 5个产品的大量评论</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b2458040bd54e9b703cbfb6f"><li>3个产品的新的数据集</li></ol><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-206d0968b24580feba4af3c003ef3e7b" data-id="206d0968b24580feba4af3c003ef3e7b"><span><div id="206d0968b24580feba4af3c003ef3e7b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580feba4af3c003ef3e7b" title="RELATED WORK"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">RELATED WORK</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245807481aaceaa4aa70470" data-id="206d0968b245807481aaceaa4aa70470"><span><div id="206d0968b245807481aaceaa4aa70470" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245807481aaceaa4aa70470" title="两个研究方向"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">两个研究方向</span></span></h3><div class="notion-text notion-block-206d0968b24580baa2b0d5afeec5091a">sentiment classification 情感分类：特定文本的情感分类——文档级、语句级</div><div class="notion-text notion-block-206d0968b245803aacbec8690dfeab9d">feature-based opinion mining 基于特征的语义挖掘</div><div class="notion-text notion-block-206d0968b24580ca94f7cc21134b2d07">&gt;更细</div><blockquote class="notion-quote notion-block-206d0968b245800b9f02d06a212c6820"><div>e.g.,“the voice quality of this phone is great and so is the reception,but the battery life is short.”</div><div class="notion-text notion-block-206d0968b24580449de8cf78e29ba71f"><em>“voice quality” ,  “reception” and“battery life” are features. The opinion on “voice quality”,“reception” are positive, and the opinion on “battery life” isnegative.</em></div></blockquote><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245806698b6d0980e699c29" data-id="206d0968b245806698b6d0980e699c29"><span><div id="206d0968b245806698b6d0980e699c29" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245806698b6d0980e699c29" title="两种基于单词短语分类的方法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>两种基于单词短语分类的方法</b></span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b24580bab449f59f0fe65dce"><li>Corpus-based approaches：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="http://t.csdnimg.cn/4EDx0">co-occurrence patterns of words </a>【10,32,34】</li></ol><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-206d0968b24580f899e0e91adc74f33a"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://uzqh6yvf0g.feishu.cn/space/api/box/stream/download/asynccode/?code=N2FjNTE1ZTc0ODEzODgxM2FjMjFjMWQ2Mjg4ZjRjZGVfU0Vlck9zSnV1MVJsbEoyeVRuclcwa3ZPVU9mRnhjZ2JfVG9rZW46WENkRGJMQXFQb2pOTER4MnFXWGNHRDhBbnRoXzE3NDg4NDYzMzQ6MTc0ODg0OTkzNF9WNA&amp;spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661&amp;t=206d0968-b245-80f8-99e0-e91adc74f33a" alt="notion image" loading="lazy" decoding="async"/></div></figure><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b2458049a4d1e63a918c88f4"><li>Dictionary-based approaches： approaches use synonyms（同义词） and antonyms（反义词） in WordNet to determine word sentiments based on a set of seed opinion words. 【1,8,13,17】</li></ol><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245802186e4dd3845288332" data-id="206d0968b245802186e4dd3845288332"><span><div id="206d0968b245802186e4dd3845288332" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245802186e4dd3845288332" title="lexicon-based method（基于词库）方法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">lexicon-based method（基于词库）方法</span></span></h3><div class="notion-text notion-block-206d0968b245809398fbebf343a07f68">【13】提出，also used in【17】，improved in 【28】by a more sophisticated method based on relaxation labeling（<a target="_blank" rel="noopener noreferrer" class="notion-link" href="http://t.csdnimg.cn/wGvjR">松弛标记法</a>）</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b2458021bfbce7eeb4e48e4f" data-id="206d0968b2458021bfbce7eeb4e48e4f"><span><div id="206d0968b2458021bfbce7eeb4e48e4f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458021bfbce7eeb4e48e4f" title="A domain specific system"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">A domain specific system</span></span></h3><div class="notion-text notion-block-206d0968b245804d87f1edbd088f9cb0">【37】分析电影评论</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245801baf26d6dcc3464f5e" data-id="206d0968b245801baf26d6dcc3464f5e"><span><div id="206d0968b245801baf26d6dcc3464f5e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245801baf26d6dcc3464f5e" title="The extraction of comparative sentences and relations"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">The extraction of comparative sentences and relations</span></span></h3><div class="notion-text notion-block-206d0968b2458048be2ae3b831d4d999">【14】比较关系抽取</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b24580c29ac5c0c16add1008" data-id="206d0968b24580c29ac5c0c16add1008"><span><div id="206d0968b24580c29ac5c0c16add1008" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580c29ac5c0c16add1008" title="相似研究"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">相似研究</span></span></h3><div class="notion-text notion-block-206d0968b24580ab80e9c76efbc12efb">Holistic lexicon-based approach整体的基于词典的方法</div><div class="notion-text notion-block-206d0968b24580c597b7f01649644ff7">识别domain opinion words 【11,16】 use conjunction rules（连接规则）：两个词使用and连接，情感倾向相同</div><blockquote class="notion-quote notion-block-206d0968b24580029871c2b97a889549"><div>“this room is beautiful and spacious”, both “beautiful” and “spacious”are positive opinion words</div></blockquote><div class="notion-text notion-block-206d0968b245807b9cc9c8e3adb41163"><b>区别：</b>虽然也使用了语言规则或习惯，但是仍有不同</div><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b2458081bb25e74e8eb83c96"><li>同一领域的相同词汇也可能表达不同的情感，要同时关注特征和opinion word</li><ol class="notion-list notion-list-numbered notion-block-206d0968b2458081bb25e74e8eb83c96"><blockquote class="notion-quote notion-block-206d0968b24580ba899ce919966a2ee7"><div>For example, in the following review sentences in thecamera domain,“the battery life is very long” and “it takes a longtime to focus”, “long”is positive in the first sentence, but negative in the second.</div></blockquote></ol></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b24580d288f9c0eb0bb772ba"><li>More flexible：不需要前置训练知识，可以online做决策</li></ol><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-206d0968b245803499e3fbfc2f3a1d0a" data-id="206d0968b245803499e3fbfc2f3a1d0a"><span><div id="206d0968b245803499e3fbfc2f3a1d0a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245803499e3fbfc2f3a1d0a" title="PROBLEM DEFINITION"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">PROBLEM DEFINITION</span></span></h2><div class="notion-text notion-block-206d0968b245802481c0dcea6c3951b8"><b>Object: </b>An object O is an entity which can be a product, person, event, organization, or topic. It is associated with a pair, O:(T, A), where T is a hierarchy or taxonomy of components (or parts), subcomponents, and soon, and A is aset of attributes of O. Each component has its own set of subcomponents and attributes.</div><div class="notion-text notion-block-206d0968b24580e398edc52277f583c0">T：组件、子组件，A：事物属性的集合</div><div class="notion-text notion-block-206d0968b2458063869affbd21a7e543">This content is only supported in a Feishu Docs</div><div class="notion-text notion-block-206d0968b24580088216c819f2bc3994">可以构成树结构，根节点为object，非根节点为component and subcomponent，节点都有对应的属性。</div><div class="notion-text notion-block-206d0968b2458083b5c1fdc9f32173d9"><b>显式特征、隐式特征：</b>句子中提到了的特征为显式特征，句子中没提到的但能推断出来的特征为隐式特征</div><blockquote class="notion-quote notion-block-206d0968b24580eca5c3f49c3b5ecfb6"><div>显式特征：“The battery life of this camera is too short”</div><div class="notion-text notion-block-206d0968b24580e58d23d076bc37a897">隐式特征：“This camera is too large”、</div></blockquote><div class="notion-text notion-block-206d0968b24580739515ccc6fde85530"><b>Opinion passage on a feature：</b>可能一段表达一个特征的观点，也可能一句话有很多特征的观点</div><div class="notion-text notion-block-206d0968b245803baf80d72689b6ed9f"><b>显式意见、隐式意见（opinion）：</b>直接表达和推断的区别</div><blockquote class="notion-quote notion-block-206d0968b245800081e3ea8565f3052c"><div>Explicit opinion:“The picture quality of this camera is amazing&quot;</div><div class="notion-text notion-block-206d0968b24580a0bf6fe721353aff97">Implicit opinion:“The earphone broke in two days&quot;</div></blockquote><div class="notion-text notion-block-206d0968b24580d7bf2ce1fb757c46e5"><b>Opinion holder：</b>观点持有者</div><blockquote class="notion-quote notion-block-206d0968b24580ac8848dce4449e4b5a"><div>“John expressed hisdisagreement on the treaty”</div></blockquote><div class="notion-text notion-block-206d0968b24580108109e6bcc3142b20"><b>Semantic orientation of an opinion</b></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b24580b6adc4d83e8489598a" data-id="206d0968b24580b6adc4d83e8489598a"><span><div id="206d0968b24580b6adc4d83e8489598a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580b6adc4d83e8489598a" title="Model definition"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Model definition</span></span></h4><div class="notion-text notion-block-206d0968b24580c984a2ced2f3eca9fe">$$F=\{f_1,f_2,\dots,f_n\}$$ 特征集合</div><div class="notion-text notion-block-206d0968b24580ba87bdc64a9681d0e6">$$W=\{w_1,w_2,\dots,w_n\}$$相关n个特征的同义词集合的集合</div><div class="notion-text notion-block-206d0968b24580ae887cf2ce41742c2a">每个opinion holder <em>j</em> 的评论在一个子集 $$S_j\subseteq F$$</div><div class="notion-text notion-block-206d0968b24580598aecd4dc70db55a7">For each feature $$f_k\in S_j$$ that opinion holder j comments on, he/she chooses aword or phrase from $$W_k$$ to  describe the feature, and then expresses a positive, negative or neutral opinion on it.</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b24580c18612fd71c5eb0b47" data-id="206d0968b24580c18612fd71c5eb0b47"><span><div id="206d0968b24580c18612fd71c5eb0b47" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580c18612fd71c5eb0b47" title="研究内容"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">研究内容</span></span></h4><div class="notion-text notion-block-206d0968b245801198cac478b9358dba">input：reviews $$D$$</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b2458052ad73cdf8b6d9f697" data-id="206d0968b2458052ad73cdf8b6d9f697"><span><div id="206d0968b2458052ad73cdf8b6d9f697" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458052ad73cdf8b6d9f697" title="$$F and W$$未知"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">$$F and W$$未知</span></span></h4><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b24580339fc6eb71538e62ee"><li>鉴别提取物品特征</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b24580a1a9dcfa765a2f12b0"><li>判断特征的情感倾向</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-206d0968b24580b8bc59e70e9338e57a"><li>将特征的同义词分组</li></ol><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b245806f952ed751dc3ca142" data-id="206d0968b245806f952ed751dc3ca142"><span><div id="206d0968b245806f952ed751dc3ca142" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245806f952ed751dc3ca142" title="$$F
$$已知$$W$$未知"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">$$F
$$已知$$W$$未知</span></span></h4><div class="notion-text notion-block-206d0968b245803b99bdf3126163184b">任务3变成使用给定的特征集合去分组</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b2458034815be98bb2990814" data-id="206d0968b2458034815be98bb2990814"><span><div id="206d0968b2458034815be98bb2990814" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458034815be98bb2990814" title="$$F and W$$已知"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">$$F and W$$已知</span></span></h4><div class="notion-text notion-block-206d0968b24580829eb8dc34b19876cd">只用完成任务2</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-206d0968b245804981a5cc122d00b2e6" data-id="206d0968b245804981a5cc122d00b2e6"><span><div id="206d0968b245804981a5cc122d00b2e6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245804981a5cc122d00b2e6" title="具体内容"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">具体内容</span></span></h4><div class="notion-text notion-block-206d0968b245808b96f6c9309c2f82bd">手机公司分析用户对一些机型的评价，任务3已有关心的特征集合以及针对不同特征的同义词集合。</div><div class="notion-text notion-block-206d0968b245808a9f95cac71eecf402">输出：对于每一个text$$d\in D$$有一个pair$$(f, SO)$$特征及其情感或意见导向</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-206d0968b2458098b9d5fed9e675c06b" data-id="206d0968b2458098b9d5fed9e675c06b"><span><div id="206d0968b2458098b9d5fed9e675c06b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458098b9d5fed9e675c06b" title="THE PROPOSED TECHNIQUE"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">THE PROPOSED TECHNIQUE</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b24580d595e9d8ffb39df58b" data-id="206d0968b24580d595e9d8ffb39df58b"><span><div id="206d0968b24580d595e9d8ffb39df58b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580d595e9d8ffb39df58b" title="关键难点"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">关键难点</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b245800db559dd7d4d844d83"><li>如何结合多元观点词语做最终判断</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b24580f0ba71f322706c69a6"><li>如何解决领域词问题</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-206d0968b24580689a72cfa470b28b87"><li>如何解决语言结构对观点词寓意倾向的改变问题</li></ol><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b24580728543e157428f19a3" data-id="206d0968b24580728543e157428f19a3"><span><div id="206d0968b24580728543e157428f19a3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580728543e157428f19a3" title="Opinion Words, Phrase and Idioms"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Opinion Words, Phrase and Idioms</span></span></h3><div class="notion-text notion-block-206d0968b24580c09fe1cfae2fdd0b9e">使用opinion lexicon （Each set is usuallyobtained through a bootstrapping process [13] using the WordNet.）</div><div class="notion-text notion-block-206d0968b24580298502d723a9f2a68e">除了原本的形容词副词，还添加了名词和动词 —— lists of context dependent opinion words</div><div class="notion-text notion-block-206d0968b245807d97e5d32b0a47a32e">POS(part-of-speech) tagging——词性</div><div class="notion-text notion-block-206d0968b2458001880efbb6345db56a"><b>idioms：</b>收集了1000多条习语</div><div class="notion-text notion-block-206d0968b2458007a920f0e36dff0a1d"><b>Non-opinion phrases containing opinion words</b></div><blockquote class="notion-quote notion-block-206d0968b2458032a7ebd79044fe7149"><div>&quot;pretty large&quot;——pretty</div></blockquote><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245804eb31be9a81f691009" data-id="206d0968b245804eb31be9a81f691009"><span><div id="206d0968b245804eb31be9a81f691009" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245804eb31be9a81f691009" title="Aggregating Opinions for a Feature"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Aggregating Opinions for a Feature</span></span></h3><div class="notion-text notion-block-206d0968b24580589b99e50514019103">This content is only supported in a Feishu Docs</div><div class="notion-text notion-block-206d0968b24580f6b5d2e1bfae059019">$$score(f)=\sum_{w_j:w_i\in s\wedge w_i\in V}{\frac{W_I.SO}{dis(w_j,f)}}$$</div><div class="notion-text notion-block-206d0968b24580d4bec8f6e6d897fe16">一些特征本身就是一个opinion word，$$score(f)$$就直接取决于该特征的情感</div><blockquote class="notion-quote notion-block-206d0968b2458029a23cd82073a5831a"><div>“This camera is very reliable&quot;</div></blockquote><div class="notion-text notion-block-206d0968b245806bbf9efef4b5936604"><b>否定规则：</b>否定词和短语与句子中表达的观点相反</div><div class="notion-text notion-block-206d0968b2458086a1d8c9a9ba2a64b5">否定词：“no”</div><div class="notion-text notion-block-206d0968b245809ba17cfd5b9bac2d77">否定模式：stop doing quit doing</div><div class="notion-text notion-block-206d0968b245808aa503fc871a72fff3">否定规则</div><blockquote class="notion-quote notion-block-206d0968b2458003b128e2ccc75390aa"><div>NN-&gt;P 没问题</div><div class="notion-text notion-block-206d0968b24580c8b4fcfb9a82b0ba49">NP-&gt;N 不好</div><div class="notion-text notion-block-206d0968b24580b9ad2ed2f61297dcc1">NNeutral-&gt;N 不起作用</div></blockquote><div class="notion-text notion-block-206d0968b2458075b441f55e395aa681"><b>包含否定词的非否定词</b></div><blockquote class="notion-quote notion-block-206d0968b24580a8b335cad3d111689e"><div>Not just</div></blockquote><div class="notion-text notion-block-206d0968b24580c9957dd377b284cc95"><b>转折从句规则</b></div><div class="notion-text notion-block-206d0968b245801e846bc3935990c42f">找不到情感倾向就在but前找，在做否定即可</div><div class="notion-text notion-block-206d0968b24580f7a926d2d1e341b69c"><b>不是but从句但是包含but词</b></div><blockquote class="notion-quote notion-block-206d0968b24580da93cedc291de3f8ad"><div>Not only but also</div></blockquote><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-206d0968b245804084d3c27d9bf7a827" data-id="206d0968b245804084d3c27d9bf7a827"><span><div id="206d0968b245804084d3c27d9bf7a827" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b245804084d3c27d9bf7a827" title="Handling Context Dependent Opinions"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Handling Context Dependent Opinions</span></span></h3><div class="notion-text notion-block-206d0968b2458060b10ff6945b17baae">整体分析法</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-206d0968b2458038ac71dfb765434646" data-id="206d0968b2458038ac71dfb765434646"><span><div id="206d0968b2458038ac71dfb765434646" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b2458038ac71dfb765434646" title="句内连词规则（Intra-sentence conjunctiong rule)"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">句内连词规则（Intra-sentence conjunctiong rule)</span></span></h4><blockquote class="notion-quote notion-block-206d0968b245806ea7a5f2f5833ef70c"><div>The battery life is very long</div><div class="notion-text notion-block-206d0968b24580cf8fdad6c1e30f85db">This camera takes great pictures and has a long battery life 运用该规则推断long是positive 因为没有but</div></blockquote><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-206d0968b24580bdbc06ea423eeba10c" data-id="206d0968b24580bdbc06ea423eeba10c"><span><div id="206d0968b24580bdbc06ea423eeba10c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580bdbc06ea423eeba10c" title="伪句内连词规则（Pseudo intra-sentence conjunction rule）"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">伪句内连词规则（Pseudo intra-sentence conjunction rule）</span></span></h4><blockquote class="notion-quote notion-block-206d0968b2458083b296e78e6559bbeb"><div>The battery life is very long</div><div class="notion-text notion-block-206d0968b2458011ae20d5ef7ab95157">The camera has a long battery life, which is great 没有and连接</div></blockquote><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-206d0968b24580fd818fffd68891b6b5" data-id="206d0968b24580fd818fffd68891b6b5"><span><div id="206d0968b24580fd818fffd68891b6b5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580fd818fffd68891b6b5" title="句间连词规则（Inter-sentence conjunction rule）"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>句间连词规则（Inter-sentence conjunction rule）</b></span></span></h4><blockquote class="notion-quote notion-block-206d0968b24580b6adc0f52fbfd74af3"><div>The picture quality is amazing. However, the battery life is short</div></blockquote><div class="notion-text notion-block-206d0968b24580d18282cf1d8069ed44"><b>Majority view</b></div><div class="notion-text notion-block-206d0968b2458077b406e910a9694389"><b>同义词反义词规则</b></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-206d0968b24580ddaf2fe06e0b06d9e8"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://uzqh6yvf0g.feishu.cn/space/api/box/stream/download/asynccode/?code=NDc4MzMyZDE1NWQwMzI0NTNhNTQwNWI3ZGQ1MGViZjhfWWhicW9VNWtDd1gyZ0VERVpzNWFlZ2hBVmpzVnc3Q1BfVG9rZW46WjN3RmJkaG9wb3lGZTd4ME1wT2NQQldjbndwXzE3NDg4NDYzMzQ6MTc0ODg0OTkzNF9WNA&amp;spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661&amp;t=206d0968-b245-80dd-af2f-e06e0b06d9e8" alt="notion image" loading="lazy" decoding="async"/></div></figure><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-206d0968b24580599fc1fa67f16a290d" data-id="206d0968b24580599fc1fa67f16a290d"><span><div id="206d0968b24580599fc1fa67f16a290d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#206d0968b24580599fc1fa67f16a290d" title="EMPIRICAL EVALUATION"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">EMPIRICAL EVALUATION</span></span></h2><div class="notion-text notion-block-206d0968b245808193c4f823285af6ad">实验数据</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-206d0968b2458029a2fff7a18c821bb2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://uzqh6yvf0g.feishu.cn/space/api/box/stream/download/asynccode/?code=OWJiYWVkZjU3ZDlkMzNhNmVkMjY2ODhjYTg5MGMwZDlfb2NNQ2FoYzdxZWNrcmlPY1dJTTVUQzAxT1Y0SG1UZGNfVG9rZW46T2NIUGJIemdUb1BaY1F4QWl1TmNmTTB5bnluXzE3NDg4NDYzMzQ6MTc0ODg0OTkzNF9WNA&amp;spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661&amp;t=206d0968-b245-8029-a2ff-f7a18c821bb2" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-206d0968b245806d992ae5b942afbaff">实验结果</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-206d0968b24580c483f0f199cbd22f79"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://uzqh6yvf0g.feishu.cn/space/api/box/stream/download/asynccode/?code=NWI0N2Y0MTA2NTc5ZGEzMzc2ODczYzVhMzUxZmU4NTBfUk5IQWdMYUI5UEdlR25KVTd2MXFWZXQzZG14RHNuM1pfVG9rZW46RlhlbmJPVFlsb3duSlZ4N09NaGNOeWgwblFmXzE3NDg4NDYzMzQ6MTc0ODg0OTkzNF9WNA&amp;spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661&amp;t=206d0968-b245-80c4-83f0-f199cbd22f79" alt="notion image" loading="lazy" decoding="async"/></div></figure></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[ACL2025]]></title>
            <link>http://preview.tangly1024.com/article/ACL2025</link>
            <guid>http://preview.tangly1024.com/article/ACL2025</guid>
            <pubDate>Wed, 04 Jun 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[绘制自然语言处理前沿图谱：ACL 2025 主要会议论文专题分析]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-208d0968b2458067aebfd98b1150c2d8"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-208d0968b24580c9b403ce94de58f0cc" data-id="208d0968b24580c9b403ce94de58f0cc"><span><div id="208d0968b24580c9b403ce94de58f0cc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580c9b403ce94de58f0cc" title="绘制自然语言处理前沿图谱：ACL 2025 主要会议论文专题分析"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://yummytanmo.github.io/ACL2025.html">绘制自然语言处理前沿图谱：ACL 2025 主要会议论文专题分析</a></b></span></span></h2><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-208d0968b2458015b32ffde5ed721e60"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:451px"><iframe class="notion-asset-object-fit" src="https://yummytanmo.github.io/ACL2025.html?spaceId=ea2f92d9-5388-4905-bb31-94997c8c0661" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b24580e1be61ddcfafc83d8b" data-id="208d0968b24580e1be61ddcfafc83d8b"><span><div id="208d0968b24580e1be61ddcfafc83d8b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580e1be61ddcfafc83d8b" title="1. 执行摘要：ACL 2025 主要会议论文核心发现"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>1. 执行摘要：ACL 2025 主要会议论文核心发现</b></span></span></h3><div class="notion-text notion-block-208d0968b24580fd8792d144ca4d0077">本报告对2025年计算语言学协会（ACL）年会主要会议录用的论文进行了全面分析，旨在揭示自然语言处理（NLP）领域的最新研究动态和未来发展趋势。分析结果表明，大型语言模型（LLMs）依然是整个领域的核心驱动力，其影响渗透到几乎所有的子领域。研究重点高度集中在LLM的能力增强、评测基准构建、效率提升、伦理考量以及多模态和智能体系统的探索上。对评测和基准测试的强烈关注，以及对效率、伦理问题和更广泛多语言能力的持续追求，共同构成了ACL 2025的研究图景。</div><div class="notion-text notion-block-208d0968b2458047b7d9c51ce88eda0e">数据显示，与LLM直接相关的研究占据了论文总数的绝大部分。具体而言，LLM核心能力（如推理、生成、长文本处理）、LLM评测与基准、LLM效率与架构以及LLM伦理AI等方向的论文数量尤为突出。同时，多模态NLP、信息获取与知识融合（特别是检索增强生成，RAG）、多语言NLP以及面向特定领域的NLP应用也展现出强劲的研究势头。</div><div class="notion-text notion-block-208d0968b24580209966e95677f2f786">ACL 2025的论文格局反映出一个领域在围绕LLM范式进行整合的同时，也在向高度专业化的研究方向拓展，以期推动LLM能力的边界、弥补其固有缺陷，并探索全新的应用前沿。这一趋势表明，NLP领域正进入一个成熟与分化并存的阶段：一方面，LLM作为基础技术的地位得到巩固；另一方面，研究者们正致力于解决LLM带来的具体挑战，并将其应用于更广泛的场景中，从而催生出大量细分且深入的研究课题。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b245809ba342f429d70bbb4e" data-id="208d0968b245809ba342f429d70bbb4e"><span><div id="208d0968b245809ba342f429d70bbb4e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b245809ba342f429d70bbb4e" title="2. 引言：勾勒ACL 2025研究版图"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>2. 引言：勾勒ACL 2025研究版图</b></span></span></h3><div class="notion-text notion-block-208d0968b24580919c99c7101bf857e9">计算语言学协会（ACL）年会是自然语言处理（NLP）领域最具影响力的顶级国际会议之一，汇集了全球研究者的最新成果，是洞察该领域发展趋势的重要窗口。本报告旨在通过对ACL 2025主要会议录用论文的系统性分析，梳理当前NLP研究的热点方向、关键挑战及未来趋势。</div><div class="notion-text notion-block-208d0968b24580238637ca27404614e1">本报告的分析方法主要基于对ACL 2025官方网站公布的主要会议论文列表（包括论文标题和作者）的细致审查 <b>1</b>。通过关键词分析、主题归类以及结合当前NLP领域的术语体系，我们对所有论文进行了系统的分类和统计。需要指出的是，由于分析仅基于论文标题和作者，可能无法完全捕捉每篇论文的全部研究内容和细微差别，这是本报告的一个局限性。</div><div class="notion-text notion-block-208d0968b24580958329d93e19d385aa">本报告后续章节将首先对ACL 2025论文进行主题分类和量化统计，然后深入剖析各个主要的研究方向，包括大型语言模型（LLMs）的持续演进、多模态NLP的兴起、信息获取与知识融合技术的发展、多语言处理的深化以及特定领域NLP应用的拓展。最后，报告将总结跨学科趋势并展望NLP领域的未来发展。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b24580ffabf9e42fbd0c9794" data-id="208d0968b24580ffabf9e42fbd0c9794"><span><div id="208d0968b24580ffabf9e42fbd0c9794" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580ffabf9e42fbd0c9794" title="3. ACL 2025 论文的主题分类与量化细分"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>3. ACL 2025 论文的主题分类与量化细分</b></span></span></h3><div class="notion-text notion-block-208d0968b24580cf8f85f324ac108d74">为了系统地把握ACL 2025的研究全貌，本报告首先对所有主要会议论文进行了主题分类。基于对论文标题的细致研读和当前NLP领域的研究热点，我们将论文归纳为十个主要研究类别。这些类别既涵盖了如大型语言模型这样的核心技术，也包括了多模态、多语言以及特定应用领域等重要方向。</div><div class="notion-text notion-block-208d0968b24580c890d7c0f2f13960b3">下表展示了ACL 2025主要会议论文在各个研究主题上的分布情况：</div><div class="notion-text notion-block-208d0968b2458010845dc3de1080cfe9"><b>表1：ACL 2025 主要会议论文按主要研究主题分布</b></div><table class="notion-simple-table notion-block-208d0968b24580d89019c6bbeaaf185d"><tbody><tr class="notion-simple-table-row notion-block-208d0968b245804b956beb128fabd3bc"><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>研究主题/类别</b></div></td><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>论文数量</b></div></td><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>占总数百分比</b></div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b2458014a33af83387d864b9"><td class="" style="width:120px"><div class="notion-simple-table-cell">1. LLM - 核心能力 (推理、生成、长文本、交互、提示、微调)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">50</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">16.67%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b245802c8405de38f4697978"><td class="" style="width:120px"><div class="notion-simple-table-cell">2. LLM - 评测、基准、鲁棒性与幻觉</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">42</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">14.00%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580218d27e71bd118c903"><td class="" style="width:120px"><div class="notion-simple-table-cell">3. LLM - 效率、扩展与架构</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">28</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">9.33%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580ae87c4cf9e28efe6ff"><td class="" style="width:120px"><div class="notion-simple-table-cell">4. LLM - 伦理AI (偏见、公平、安全、可解释性、隐私、虚假信息)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">33</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">11.00%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580acbbb5c1cf93dc8d31"><td class="" style="width:120px"><div class="notion-simple-table-cell">5. 多模态NLP (视觉、音频、视频、数据集、交互)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">35</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">11.67%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580f096bff09981b79a60"><td class="" style="width:120px"><div class="notion-simple-table-cell">6. 信息获取与知识融合 (RAG、问答、知识图谱、事实核查)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">27</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">9.00%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580b2b21efc077e9ac992"><td class="" style="width:120px"><div class="notion-simple-table-cell">7. 多语言与跨语言NLP (模型、低资源、机器翻译、文化适应)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">23</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">7.67%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b2458004bfb6c9b334c37118"><td class="" style="width:120px"><div class="notion-simple-table-cell">8. 专业领域NLP应用 (医疗、法律、电商、代码、社科、教育、金融)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">38</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">12.67%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b245806c85e8c25e948177ec"><td class="" style="width:120px"><div class="notion-simple-table-cell">9. 基础NLP任务与技术 (对话、情感、GEC、信息抽取、语言学 - 演进中)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">15</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">5.00%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b24580c18c50e416b74f640d"><td class="" style="width:120px"><div class="notion-simple-table-cell">10. NLP研究生态与方法论 (综述、通用数据集/评测、人机交互)</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">9</div></td><td class="" style="width:120px"><div class="notion-simple-table-cell">3.00%</div></td></tr><tr class="notion-simple-table-row notion-block-208d0968b2458010a339f92919f103f4"><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>总计</b></div></td><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>300</b></div></td><td class="" style="width:120px"><div class="notion-simple-table-cell"><b>100.00%</b></div></td></tr></tbody></table><div class="notion-text notion-block-208d0968b24580fe9f4fff64c27ee100"><em>注：论文分类基于其主要研究贡献，部分论文可能涉及多个主题。此处的计数反映了其核心归属。数据来源于对 </em><em><b>1</b></em><em> 所列论文标题的分析。</em></div><div class="notion-text notion-block-208d0968b24580e4be18f36245ed5c82">从表1的初步观察可见，大型语言模型（LLM）相关研究占据了主导地位。将前四个类别（LLM核心能力、LLM评测、LLM效率、LLM伦理AI）相加，直接与LLM相关的论文数量达到了153篇，超过了总论文数的一半（51%）。这清晰地表明LLM不仅是NLP领域的一个重要分支，更是当前研究的核心引擎和平台。此外，多模态NLP（11.67%）和面向特定领域的NLP应用（12.67%）也显示出较高的研究热度，反映了NLP技术向更复杂场景和实际应用拓展的趋势。信息获取与知识融合（9.00%）以及多语言NLP（7.67%）同样是重要的研究方向。相对而言，传统的基础NLP任务虽然仍在发展，但其论文占比较小（5.00%），这可能意味着许多此类任务正被整合到LLM的框架下进行研究。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b24580eaacb7c8e853bcef64" data-id="208d0968b24580eaacb7c8e853bcef64"><span><div id="208d0968b24580eaacb7c8e853bcef64" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580eaacb7c8e853bcef64" title="4. 主要和新兴研究方向深度剖析"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4. 主要和新兴研究方向深度剖析</b></span></span></h3><div class="notion-text notion-block-208d0968b2458001bcebd233f02de6bb">本节将对ACL 2025论文所反映出的主要研究方向进行更深入的定性分析，并通过列举具体的论文标题 <b>1</b> 作为例证，揭示各领域的关键进展和潜在趋势。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b245804e893ee5fcc393531c" data-id="208d0968b245804e893ee5fcc393531c"><span><div id="208d0968b245804e893ee5fcc393531c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b245804e893ee5fcc393531c" title="4.1. 大型语言模型 (LLMs)：持续主导与多元发展"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.1. 大型语言模型 (LLMs)：持续主导与多元发展</b></span></span></h4><div class="notion-text notion-block-208d0968b2458029baeac572c9180381">大型语言模型无疑是ACL 2025中最耀眼的明星，其影响力贯穿了几乎所有NLP研究分支。LLM不再仅仅是一个独立的研究课题，而已然演化为驱动当前多数NLP创新的基础性技术平台。对LLM的研究并非单一维度，而是呈现出一个复杂且多方面的生态系统，涵盖了从提升核心能力、理解内在机理、克服固有局限，到确保其安全负责任应用等各个层面。</div><div class="notion-text notion-block-208d0968b24580f2929fc79868dd430c">这种多方位探索的格局，体现在大量论文分别聚焦于LLM的不同侧面。例如，一些研究致力于增强模型的推理能力（如《Capture the Key in Reasoning to Enhance CoT Distillation Generalization》），另一些则专注于提升生成质量（如《Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models》）。同时，评测（如《JuStRank: Benchmarking LLM Judges for System Ranking》）、效率（如《TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs》）以及伦理（如《The Impossibility of Fair LLMs》）等方面的研究也层出不穷 <b>1</b>。这些看似分散的研究点，实际上共同构成了一个旨在全面推进LLM技术的宏大研究议程。一个子领域的进展，往往能为其他子领域带来新的可能性或提出新的要求，整个LLM研究领域正是在这种相互促进、协同演化的动态中不断前进。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580aaa33fd260fbc0e22c" data-id="208d0968b24580aaa33fd260fbc0e22c"><span><div id="208d0968b24580aaa33fd260fbc0e22c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580aaa33fd260fbc0e22c" title="4.1.1. LLM核心能力：推理、生成、长文本与交互"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.1.1. LLM核心能力：推理、生成、长文本与交互</b></span></span></h4><div class="notion-text notion-block-208d0968b24580f69d46e54ab59f736a">研究者们在提升LLM的核心智能方面投入了巨大精力，特别是在推理、可控生成、长文本理解以及交互式应用等关键能力上。</div><ul class="notion-list notion-list-disc notion-block-208d0968b2458062b438d768e809f3b1"><li><b>推理能力</b>：增强LLM的逻辑推导、数学解题和链式思维（Chain-of-Thought, CoT）能力是本届会议的一大热点。例如，《Capture the Key in Reasoning to Enhance CoT Distillation Generalization》探索了如何提炼和泛化CoT推理过程，《ProcessBench: Identifying Process Errors in Mathematical Reasoning》则关注数学推理中的错误识别 <b>1</b>。更有研究如《Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework》和《MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset》致力于构建更完备的逻辑推理框架和形而上推理能力的基准测试 <b>1</b>。这些工作表明，学术界正努力推动LLM从简单的模式匹配向更鲁棒、可验证的复杂推理能力迈进，这通常涉及到多步骤、结构化的思考过程。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245809ab5fefbb596459b1a"><li><b>生成能力</b>：在文本生成方面，研究重点在于提升生成内容的可控性、多样性和特定任务的适用性。代码生成是一个显著的例子，如《Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models》通过树状指令进化来优化代码生成 <b>1</b>。此外，《Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models》探索了利用扩散模型进行可控长文本生成，《Generating Diverse Training Samples for Relation Extraction with Large Language Models》则利用LLM进行数据增强以服务于关系抽取等下游任务 <b>1</b>。这些研究体现了对生成实用、高保真输出的追求。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458042bec2cd48486c650e"><li><b>长文本处理</b>：扩展LLM的有效上下文窗口是另一个关键研究方向，对于文档理解、长篇摘要和复杂对话等任务至关重要。相关工作包括《Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models》、《Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free Method》以及《Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models》<b>1</b>。基准测试如《LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks》也应运而生，旨在更全面地评估长文本处理能力 <b>1</b>。长文本处理能力的瓶颈显而易见，解决方案从模型架构调整到数据策略优化不一而足，足见其重要性。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580cda7e7c77cddda11ce"><li><b>交互与智能体</b>：将LLM发展为能够与环境交互、使用工具并扮演智能体角色的系统，是NLP领域一个令人兴奋的新方向。例如，《MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments》研究了LLM在复杂社交环境中的表现，《CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System》和《AndroidLab: Developing and Evaluating Android Agents in A Reproducible Environment》则分别探索了LLM在代码编译和安卓环境控制方面的智能体应用 <b>1</b>。这标志着LLM正从单纯的文本处理器向能够执行目标导向行为的实体转变。</li></ul><div class="notion-text notion-block-208d0968b2458089b40bffb5baacbcc4">这些核心能力的提升并非孤立进行，而是相互依存、相互促进的。例如，有效的长文本处理能力是LLM对大型文档进行复杂推理的基础；而高质量的生成能力则是智能体系统清晰传达其行动或发现的前提。一篇名为《LongBench v2》的论文 <b>1</b> 明确地将其目标设定为“对现实长文本多任务的更深层次理解和推理”，直接将长文本处理与推理能力联系起来。同样，像《CompileAgent》<b>1</b> 这样的工作，需要LLM理解指令（推理）、生成命令（生成）并处理输出（通常是长而复杂的文本）。因此，这些能力的协同发展对LLM的整体进步至关重要，某一方面的局限（如上下文长度不足）可能会制约其他方面（如智能体的复杂多跳推理）的发展。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580949cc3def8e7ca6656" data-id="208d0968b24580949cc3def8e7ca6656"><span><div id="208d0968b24580949cc3def8e7ca6656" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580949cc3def8e7ca6656" title="4.1.2. LLM的评测、基准与鲁棒性"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.1.2. LLM的评测、基准与鲁棒性</b></span></span></h4><div class="notion-text notion-block-208d0968b245809cb412e793d93d7f37">随着LLM能力的日益复杂化，如何对其进行可靠和全面的评测成为了一个核心挑战。ACL 2025涌现了大量关于新基准、新指标以及评测方法论本身的研究。</div><ul class="notion-list notion-list-disc notion-block-208d0968b24580fdb647d6ea83c26115"><li><b>评测基准与指标</b>：研究者们开发了众多新的基准测试，旨在评估LLM在各种专门任务上的表现。例如，《ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming》关注人与LLM在编程竞赛中的表现，《RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios》则侧重于规则指导下的推理能力 <b>1</b>。同时，对评测方法本身的探讨也十分引人注目，如《JuStRank: Benchmarking LLM Judges for System Ranking》研究了使用LLM作为评测者的可行性，而《A Measure of the System Dependence of Automated Metrics》则分析了自动化指标的系统依赖性 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245808eaf40e4ef190ec8f1"><li><b>鲁棒性</b>：评估和提升LLM在面对对抗性攻击、分布外输入和“陷阱问题”时的性能稳定性，是确保模型可靠性的关键。相关研究如《Wait, that’s not an option: LLMs Robustness with Incorrect Multiple-Choice Options》和《What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs》深入探讨了LLM的鲁棒性问题 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580e0959eddf36c7eeaac"><li><b>幻觉缓解</b>：LLM生成事实错误或无意义信息的“幻觉”问题，依然是研究的重中之重。大量工作致力于检测、减少和理解幻觉产生的原因。例如，《HALoGEN: Fantastic LLM Hallucinations and Where to Find Them》对幻觉现象进行了探索，《MPVStance: Mitigating Hallucinations in Stance Detection with Multi-Perspective Verification》提出了多视角验证来缓解立场检测中的幻觉，《Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence》则从视觉感知的角度研究大型视觉语言模型中的幻觉问题 <b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b245806ba2c9e7047045f766">当前，NLP领域正经历一场“评测军备竞赛”，以寻求对模型能力更深层次的理解。大量新基准的涌现（如 <b>1</b> 中的“ELABORATION”、“RuleArena”、“LongBench v2”）以及对评测方法本身的批判和改进（如 <b>1</b> 中的“JuStRank”、“A Measure of the System Dependence of Automated Metrics”、“Call for Rigor in Reporting Quality of Instruction Tuning Data”）表明，现有的评测体系正迅速变得不足或易被“应试”。社群不仅在创造更多的测试，更在反思如何进行测试。将LLM用作评测裁判（如 <b>1</b> 中的“JuStRank”）的尝试，反映了研究者在寻求超越简单指标、实现更具可扩展性和细致性的评测方法，尽管这本身也引入了新的偏见和挑战。这种趋势揭示了对模型智能进行更深入、更真实评估的渴望，而非仅仅追求在排行榜上的高分。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b245807ba9c9d186ae253afa" data-id="208d0968b245807ba9c9d186ae253afa"><span><div id="208d0968b245807ba9c9d186ae253afa" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b245807ba9c9d186ae253afa" title="4.1.3. LLM的效率、扩展与架构创新"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.1.3. LLM的效率、扩展与架构创新</b></span></span></h4><div class="notion-text notion-block-208d0968b24580b6a3bfdfba68748e16">在追求模型能力提升的同时，如何使LLM更高效、更易于部署，也是一个核心议题。</div><ul class="notion-list notion-list-disc notion-block-208d0968b2458063b6bbf3cbcefe69b4"><li><b>计算效率</b>：研究者们探索了多种技术以降低LLM的计算和存储开销，包括量化（如《L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models》和《PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models》）、剪枝（如《TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs》）、知识蒸馏以及高效注意力机制（如《KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding》）和加速推理（如《FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling》）<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245806692ded92f009fd776"><li><b>模型架构</b>：除了对现有Transformer架构进行优化，研究者们也在探索替代性架构。例如，《The Hidden Attention of Mamba Models》对基于Mamba的新型架构的内部机制进行了分析 <b>1</b>。混合专家模型（MoE）也因其高效扩展潜力而受到关注，如《Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245808293e9f2a4928dbb73"><li><b>扩展法则</b>：对LLM扩展法则（Scaling Laws）的研究仍在继续，旨在理解模型规模、数据量和计算资源之间的关系，并指导未来的模型设计和训练策略，如《P$^2$ Law: Scaling Law for Post-Training After Model Pruning》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b24580a8b6a6de551a93d548">“实用性优先”的趋势日益明显。尽管继续扩大模型规模仍是前沿探索的一部分，但一股强大的逆流正致力于使这些强大的模型能够在资源受限的环境中运行，或以更低的运营成本部署。大量关于模型效率的研究（如 <b>1</b> 中的“TrimLLM”、“L4Q”、“PTQ1.61”、“KV-Latent”）直接应对了LLM因其庞大规模和高昂计算成本而难以广泛应用、难以被小型实验室研究以及难以在端侧设备部署的挑战。效率研究的成功，如量化和新架构的提出，可以极大地拓宽LLM技术的应用范围和可及性，这可能导致模型发展的分化：一端是巨大的前沿模型，另一端则是高效、专业化的小型模型。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580d38cc3d8e5a4ac9800" data-id="208d0968b24580d38cc3d8e5a4ac9800"><span><div id="208d0968b24580d38cc3d8e5a4ac9800" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580d38cc3d8e5a4ac9800" title="4.1.4. LLM时代的伦理AI：偏见、公平、安全与可解释性"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.1.4. LLM时代的伦理AI：偏见、公平、安全与可解释性</b></span></span></h4><div class="notion-text notion-block-208d0968b24580bead2feacf2f20388d">随着LLM能力的增强及其应用的普及，伦理问题受到了前所未有的关注。</div><ul class="notion-list notion-list-disc notion-block-208d0968b245803da5f5f1831c7c77cf"><li><b>偏见与公平</b>：识别和减轻LLM在表示和输出中存在的社会偏见（如性别、种族偏见）是核心议题。相关工作包括对公平LLM不可能性的理论探讨（《The Impossibility of Fair LLMs》）、超越简单测试的偏见评估方法（《Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation》）、多语言伦理偏见的研究（《Delving into Multilingual Ethical Bias...》）以及对特定偏见现象的分析（如《On the Mutual Influence of Gender and Occupation in LLM Representations》和《White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs》）<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580d88717c0dd1842622b"><li><b>安全与保障</b>：防止LLM被滥用、抵御对抗性攻击和越狱（jailbreaking）、确保模型不生成有害内容，是安全研究的重点。例如，《Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation》评估了LLM被用于个性化虚假信息生成的风险，《Root Defense Strategies: Ensuring Safety of LLM at the Decoding Level》从解码层面探讨安全保障，《Jailbreak Large Vision-Language Models Through Multi-Modal Linkage》研究了多模态模型的越狱问题，《Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training》和《SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model》则分别关注提升LLM的拒答能力和RAG系统的安全性 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580859dd6d6914c9ad506"><li><b>可解释性与透明度</b>：理解LLM的内部工作机制并为其预测提供合理解释，对于建立信任和调试模型至关重要。相关研究如《TAGExplainer: Narrating Graph Explanations for Text-Attributed Graph Learning Models》、《ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text Classification》和《Position-aware Automatic Circuit Discovery》等 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458050a830de17ebc49cb2"><li><b>虚假信息与内容溯源</b>：检测和对抗机器生成的虚假信息，以及对生成内容进行溯源（如水印技术《Ensemble Watermarks for Large Language Models》<b>1</b>），也是重要的研究方向。例如，《Real-time Fake News from Adversarial Feedback》、《Detection of Human and Machine-Authored Fake News in Urdu》和《TripleFact: Defending Data Contamination in the Evaluation of LLM-driven Fake News Detection》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b24580ff95b8d1ecaf229b57">伦理考量正日益成为NLP研究的核心组成部分，而非次要的补充。大量关于伦理AI的研究（涵盖偏见、安全、可解释性、虚假信息等）表明，学术界正将这些问题视为核心研究挑战。过去AI系统因固化偏见或易被滥用而受到批评，NLP社群正积极主动地应对LLM可能带来的更大社会影响。诸如《The Impossibility of Fair LLMs》<b>1</b> 这样的论文表明，这些问题并非简单的技术修复所能解决，而是深刻的、有时甚至是矛盾的挑战，需要基础性研究。这意味着伦理考量正融入LLM开发的全生命周期。该领域正从仅仅识别问题转向积极开发缓解技术和框架，尽管根本性的挑战依然存在。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580adb4ffe5879f1a3f4d" data-id="208d0968b24580adb4ffe5879f1a3f4d"><span><div id="208d0968b24580adb4ffe5879f1a3f4d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580adb4ffe5879f1a3f4d" title="4.2. 多模态NLP：语言、视觉、音频及其他模态的协同"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.2. 多模态NLP：语言、视觉、音频及其他模态的协同</b></span></span></h4><div class="notion-text notion-block-208d0968b24580789916cd4d0697c912">能够处理和整合来自多种模态（文本、图像、音频、视频等）信息的模型，正成为NLP领域一个快速增长的研究方向。</div><ul class="notion-list notion-list-disc notion-block-208d0968b2458097a623cfc2cbd1c156"><li><b>视觉语言模型 (VLMs)</b>：VLMs的研究重点包括视觉问答（VQA）、图像描述生成、以及基于视觉和文本数据的复杂推理。例如，《Can Multimodal Large Language Models Understand Spatial Relations?》探讨了VLM对空间关系的理解能力，《PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension》则关注VLM对多模态幽默的理解 <b>1</b>。提升VLM推理能力的工作如《Improve Vision Language Model Chain-of-thought Reasoning》，而安全性研究则有《Jailbreak Large Vision-Language Models Through Multi-Modal Linkage》<b>1</b>。更有趣的应用如《TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding》，尝试让VLM对数学定理进行基于视频的多模态解释 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458078a60ee9332c1a8c16"><li><b>音频与语音集成</b>：语音合成、语音转换、以及在上下文中理解口语的研究也取得了显著进展。例如，《Autoregressive Speech Synthesis without Vector Quantization》提出了一种无需矢量量化的自回归语音合成方法，《Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling》专注于表现力强的零样本语音转换 <b>1</b>。此外，《In-the-wild Audio Spatialization with Flexible Text-guided Localization》研究了真实环境下的音频空间化，而《Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models》则致力于评估大型音频语言模型的开放域对话理解能力 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458056b76ed5eb56d59398"><li><b>多模态数据集与基准</b>：为了训练和评估日益复杂的多模态系统，新的数据集和基准不断涌现。例如，《LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating》提供了一个包含长文档的多模态基准，《BQA: Body Language Question Answering Dataset for Video Large Language Models》专注于身体语言问答，《Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues》则收集了大规模的包含非语言线索的视频对话数据 <b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b24580a8baf0c3c5974a3765">多模态NLP的研究趋势正从简单地并行处理不同模态信息，转向实现更深层次的模态融合、跨模态推理，并更有效地将语言在其他模态中进行“锚定”（grounding）。诸如《Can Multimodal Large Language Models Understand Spatial Relations?》<b>1</b> 和《Improve Vision Language Model Chain-of-thought Reasoning》<b>1</b> 这样的论文，其关注点在于需要视觉和文本信息紧密结合的复杂推理任务。同时，像《BQA: Body Language Question Answering Dataset》<b>1</b> 和《Speaking Beyond Language...》<b>1</b> 这样的数据集的创建，表明了对能够测试细致入微、综合理解能力的资源的需求，而非仅仅是表面的关联。早期的多模态模型可能擅长图像描述等任务，而当前的研究则致力于让模型能够解释视频为何有趣（如 <b>1</b> 中的“PunchBench”），或理解涉及视觉和文本元素的复杂指令。这要求更复杂的模态融合架构和表示学习方法，以期构建出对通过多感官感知的世界具有更整体“理解”的模型。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580309e27e0af473da316" data-id="208d0968b24580309e27e0af473da316"><span><div id="208d0968b24580309e27e0af473da316" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580309e27e0af473da316" title="4.3. 信息获取与知识整合"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.3. 信息获取与知识整合</b></span></span></h4><div class="notion-text notion-block-208d0968b24580c48b84eb637df2f4c7">如何利用LLM有效访问、检索和推理海量信息，是NLP领域的另一个核心议题。</div><ul class="notion-list notion-list-disc notion-block-208d0968b245804f9219c455fad5d852"><li><b>检索增强生成 (RAG)</b>：RAG已成为一种主流范式，用于将LLM的输出锚定在外部知识源上，以对抗幻觉、提供最新信息。ACL 2025中有大量关于RAG的研究，例如《HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases》探索了混合知识库的RAG，《DioR: Adaptive Cognitive Detection and Contextual Retrieval Optimization for Dynamic Retrieval-Augmented Generation》关注动态RAG中的自适应检索与优化，《MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation》提出了多智能体过滤的RAG框架 <b>1</b>。对RAG鲁棒性的研究如《RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation》，而《Pandora’s Box or Aladdin’s Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models》则深入分析了RAG中噪声数据的影响 <b>1</b>。这些工作表明，优化检索、生成以及两者之间的交互，包括处理嘈杂或冲突的检索信息，是RAG研究的重点。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458085bd3ef5dd876b8d50"><li><b>问答 (QA)</b>：复杂问答、多跳推理以及基于结构化和非结构化数据的问答系统依然是研究热点。例如，《ReSCORE: Label-free Iterative Retriever Training for Multi-hop Question Answering with Relevance-Consistency Supervision》提出了一种无标签迭代检索器训练方法用于多跳问答，《Doc-React: Multi-page Heterogeneous Document Question-answering》则专注于多页异构文档的问答 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580dab3b0f57aafcee4fa"><li><b>知识图谱与结构化数据</b>：利用LLM与知识图谱交互、查询知识图谱、甚至生成知识图谱，以及对表格等结构化数据进行推理，也吸引了广泛关注。相关工作如《RelationalCoder: Relational Representation of Complex Tables for Program-Based Processing and Reasoning》、《HyperFM: Fact-Centric Multimodal Fusion for Link Prediction over Hyper-Relational Knowledge Graphs》和《TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b245803b976ff244d7db14ad">LLM与外部知识源之间的关系正变得日益共生和演进。一方面，LLM越来越依赖外部知识（通过RAG）来提升其输出的事实性和相关性。另一方面，LLM本身也正成为更好地理解、处理甚至生成结构化知识的工具。大量关于RAG的论文（如 <b>1</b> 中的“HybGRAG”、“DioR”、“MAIN-RAG”）清晰地展示了用外部数据增强LLM的趋势。同时，像《RelationalCoder》<b>1</b> 这样的工作则利用LLM处理结构化数据。这是一个双向的过程：LLM消费外部知识以改进自身，同时LLM也生产或处理结构化知识。然而，这种整合并非没有挑战，正如《Pandora’s Box or Aladdin’s Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise...》<b>1</b> 所指出的，RAG的引入可能带来新的问题，例如检索到的噪声数据的影响。未来的研究可能会聚焦于更复杂的RAG技术，包括更好的检索器、评估检索信息质量的机制，以及更鲁棒的内部（参数化）知识和外部（非参数化）知识的融合方法。 “知晓”（参数化知识）和“获取知识”（RAG）之间的界限将持续模糊。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580a99b71e7651c19f6b1" data-id="208d0968b24580a99b71e7651c19f6b1"><span><div id="208d0968b24580a99b71e7651c19f6b1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580a99b71e7651c19f6b1" title="4.4. 多语言与跨语言应用"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.4. 多语言与跨语言应用</b></span></span></h4><div class="notion-text notion-block-208d0968b24580ce83a3da370a638b8d">解决全球语言多样性带来的挑战，是NLP领域持续努力的方向。</div><ul class="notion-list notion-list-disc notion-block-208d0968b245809391c4d1a21f19a286"><li><b>多语言模型与基准</b>：开发和评估能够在多种语言上执行任务的模型是核心工作。例如，《M-RewardBench: Evaluating Reward Models in Multilingual Settings》评估了多语言环境下的奖励模型，《BelarusianGLUE: Towards a Natural Language Understanding Benchmark for Belarusian》为白俄罗斯语构建了NLU基准 <b>1</b>。提升多语言模型自然度的研究如《Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs》，而《MUSTS: MUltilingual Semantic Textual Similarity Benchmark》则提供了多语言语义相似度基准 <b>1</b>。预训练方面的研究有《LangSAMP: Language-Script Aware Multilingual Pretraining》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580ef863dee538da42c35"><li><b>低资源语言</b>：为数据有限的语言构建NLP工具的技术备受关注，包括创新的数据增强和迁移学习方法。例如，《Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion》研究了LLM对阿拉伯语的二语习得，《Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books》探索了利用代码增强语法书翻译极低资源语言的方法 <b>1</b>。此外，《Improving Parallel Sentence Mining for Low-Resource and Endangered Languages》和《Understanding In-context Machine Translation for Low-Resource Languages: A Case Study on Manchu》也分别关注了低资源语言的平行句挖掘和上下文学习翻译问题 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458032a481f3265b974d8c"><li><b>机器翻译 (MT)</b>：在LLM时代，机器翻译的质量、鲁棒性和评估方法持续取得进展。例如，《Did Translation Models Get More Robust Without Anyone Even Noticing?》探讨了翻译模型的鲁棒性演变，《Unveiling the Power of Source: Source-based Minimum Bayes Risk Decoding for Neural Machine Translation》则研究了基于源语言的解码策略 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580168362d1170e9c64fe"><li><b>跨语言迁移与文化适应</b>：如何将在一种语言上学到的知识应用于其他语言，以及如何使模型适应特定的文化背景，是重要的研究课题。相关工作包括《Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon》、《Cultural Learning-Based Culture Adaptation of Language Models》和《Towards Geo-Culturally Grounded LLM Generations》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b245800bb1e9c8c7997af052">多语言NLP研究正从简单的语言覆盖向更深层次的问题成熟。研究重点不再仅仅是将模型扩展到更多语言，而是转向解决真正的跨语言理解、有效处理数据稀缺性、确保文化适宜性以及评估多语言输出的细微质量等问题（例如 <b>1</b> 中的“Do Large Language Models have an English Accent?”）。论文不再仅仅是关于“某种X语言的模型”，而是关于如何更好地进行多语言NLP：例如为特定语族构建基准（如 <b>1</b> 中的“BelarusianGLUE”），为极低资源语言开发技术（如 <b>1</b> 中的“Read it in Two Steps...”），以及进行文化适应（如 <b>1</b> 中的“Cultural Learning-Based...”）。诸如《Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models》<b>1</b> 和《Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness...》<b>1</b> 等工作，指出了对多语言模型中挑战的更深入调查。这表明该领域认识到，仅仅用更多语言进行训练是不够的。在表示、迁移和评估方面存在根本性挑战，需要解决这些挑战才能实现公平有效的多语言AI。未来的多语言研究可能会更侧重于语言支持的质量而非数量，强调真正的理解、低资源场景的鲁棒性以及具有文化意识的生成。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b2458042a611ea3527453b92" data-id="208d0968b2458042a611ea3527453b92"><span><div id="208d0968b2458042a611ea3527453b92" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b2458042a611ea3527453b92" title="4.5. 专业领域NLP应用"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.5. 专业领域NLP应用</b></span></span></h4><div class="notion-text notion-block-208d0968b2458028967be29eb095d35f">将NLP技术应用于特定领域，通常需要领域知识的融合和模型的专门适配。</div><ul class="notion-list notion-list-disc notion-block-208d0968b24580f5ae40fab81e0f1022"><li><b>医疗健康</b>：自动化报告生成、临床编码、医学问答等是热点方向。例如，《The Impact of Auxiliary Patient Data on Automated Chest X-Ray Report Generation and How to Incorporate It》研究了辅助数据在胸片报告生成中的作用，《Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review》探讨了AI与临床编码流程的对齐 <b>1</b>。非洲医学问答基准《AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset》和单细胞生物学基础语言模型综述《A Survey on Foundation Language Models for Single-cell Biology》也反映了这一趋势 <b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245803fa49ff7a73f44dd8a"><li><b>法律领域</b>：基于智能体的法律任务辅助、法律文本解读、合同审查等应用受到关注。例如，《LegalAgentBench: Evaluating LLM Agents in Legal Domain》、《Automating Legal Concept Interpretation with LLMs: Retrieval, Generation, and Evaluation》以及合同自动审查条款推荐基准《ProvBench: A Benchmark of Legal Provision Recommendation for Contract Auto-Reviewing》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458090b94add290e107ab5"><li><b>电子商务与金融</b>：电商领域的脚本规划、属性挖掘、多模态检索，以及金融领域的LLM智能体决策等是新兴应用。例如，《EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association》、《Open-World Attribute Mining for E-Commerce Products with Multimodal Self-Correction Instruction Tuning》和金融决策基准《INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580489104de39df8eebab"><li><b>代码生成与理解</b>：LLM在编程任务、缺陷检测、代码简化等方面的应用持续火热。例如，《Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models》、《LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs》和《WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245800c89a9e97124144866"><li><b>社会科学与人文学科</b>：利用NLP分析社交媒体、文学文本、理解文化现象等。例如，《Literature Meets Data: A Synergistic Approach to Hypothesis Generation》、《Capturing Author Self Beliefs in Social Media Language》和《When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458098bef7d9274bf2be62"><li><b>教育领域</b>：面向教育的摘要生成、评估学生写作等。例如，《From Information to Insight: Leveraging LLMs for Open Aspect-Based Educational Summarization》和《LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b245801caddec9d2e42bccad">领域自适应是实现NLP技术真实世界影响力的关键。尽管通用LLM功能强大，但其在专业领域的有效应用需要仔细的调整、领域特定数据的微调，并且通常需要整合领域知识。大量论文聚焦于将NLP/LLM应用于医疗、法律、金融、电商等特定领域。这些领域拥有独特的术语、推理模式和数据格式，通用模型若不进行适配，性能往往不佳。领域特定基准的创建（例如 <b>1</b> 中的“LegalAgentBench”、“AfriMed-QA”、“EcomScriptBench”、“INVESTORBENCH”）突显了为特定应用定制和评估模型的公认需求。像《Aligning AI Research with the Needs of Clinical Coding Workflows》<b>1</b> 这样的论文明确讨论了通用AI能力与特定领域具体需求之间的差距。因此，大量的研究工作正致力于弥合通用LLM能力与专业领域细致需求之间的鸿沟，这不仅涉及微调，还包括开发针对这些领域的新评估方法和数据集。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b2458014ace9cbc6c1f48693" data-id="208d0968b2458014ace9cbc6c1f48693"><span><div id="208d0968b2458014ace9cbc6c1f48693" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b2458014ace9cbc6c1f48693" title="4.6. LLM时代的基础NLP任务与技术演进"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.6. LLM时代的基础NLP任务与技术演进</b></span></span></h4><div class="notion-text notion-block-208d0968b2458070aabccb909649e9d7">传统的NLP任务在LLM时代正被重新审视或进一步发展，LLM常常作为其中的一个组件或强大的基线模型。</div><ul class="notion-list notion-list-disc notion-block-208d0968b2458071ad78c5d530ebc588"><li><b>对话系统</b>：少样本意图分类、复杂对话管理、个性化智能体等是研究重点。例如，《Dynamic Label Name Refinement for Few-Shot Dialogue Intent Classification》、《Battling against Tough Resister: Strategy Planning with Non-collaborative Dialogues》和《In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458057b689d67a08ab2d61"><li><b>情感分析与情绪识别</b>：跨语言情感分析、多模态情绪检测等方向有所进展。例如，《LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation》、《Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge》和《ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in Conversation》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580deb6d2c21be4f0e1cb"><li><b>语法错误纠正 (GEC)</b>：可解释性评估是GEC领域的新关注点，如《CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245803cba42dcac4e276507"><li><b>主题建模</b>：将LLM整合到主题建模框架中，如《Neural Topic Modeling with Large Language Models in the Loop》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b245803a974af6840844eb59">基础NLP任务正被LLM重塑，而非完全取代。尽管LLM能够零样本或少样本执行许多这类任务，但专门的研究仍在继续改进方法论，通常是通过新颖的方式利用LLM（例如，用于数据增强、作为更大系统中的组件，或作为特定语言现象的研究对象）。仍有论文致力于情感分析（如 <b>1</b> 中的“LACA”）、语法错误纠正（如 <b>1</b> 中的“CLEME2.0”）和对话系统等任务。人们可能认为LLM使得对这些任务的专门研究变得过时。然而，这些主题的持续存在表明，要么LLM对于这些任务的所有细微之处尚非完美解决方案，要么LLM正作为强大的新工具被整合到这些研究领域中，而不是完全取代它们。例如，《LACA...with LLM Data Augmentation》<b>1</b> 明确展示了LLM被用于改进传统任务。而《Neural Topic Modeling with Large Language Models in the Loop》<b>1</b> 则展示了整合应用。因此，基础NLP任务正在演变。研究重点可能从从头开始构建特定任务模型，转向理解如何最好地利用或引导LLM完成这些任务，或解决LLM仍然不足的剩余差距。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-208d0968b24580b398a7c0a2e0b39840" data-id="208d0968b24580b398a7c0a2e0b39840"><span><div id="208d0968b24580b398a7c0a2e0b39840" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580b398a7c0a2e0b39840" title="4.7. NLP研究生态：数据集、基准与评测方法论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>4.7. NLP研究生态：数据集、基准与评测方法论</b></span></span></h4><div class="notion-text notion-block-208d0968b2458057accff8c55fc44c66">这是一个贯穿各研究方向的主题，但值得特别强调的是社群在构建和批判NLP研究基础设施方面所做的努力。</div><ul class="notion-list notion-list-disc notion-block-208d0968b24580f5a3d4c04b3c7ba8f3"><li><b>新数据集创建</b>：针对各种任务、语言和模态创建新的数据集资源。例如，《EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association》、《BelarusianGLUE: Towards a Natural Language Understanding Benchmark for Belarusian》、《AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset》、《LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating》以及《HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580b98ecceeaa6a05d9d2"><li><b>新颖的基准测试方法</b>：超越标准的排行榜，专注于评估特定能力或模拟真实世界场景。例如，《RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios》、《JuStRank: Benchmarking LLM Judges for System Ranking》、《LegalAgentBench: Evaluating LLM Agents in Legal Domain》和《EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents》<b>1</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b2458024bdbeeea18d513904"><li><b>评测方法的批判与改进</b>：质疑现有指标的有效性，提出新的评估框架，并呼吁更严格的评测标准。例如，《A Measure of the System Dependence of Automated Metrics》、《Call for Rigor in Reporting Quality of Instruction Tuning Data》以及一篇引人深思的论文《Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above》<b>1</b>。</li></ul><div class="notion-text notion-block-208d0968b24580f59607fb3cde4c2dd2">NLP研究社群内部存在着一种自我修正和追求严谨的文化。社群正积极参与关于如何进行和评估研究的元层面讨论。许多论文并非仅仅展示新模型或SOTA结果，而是聚焦于我们如何衡量进展（如 <b>1</b> 中的“A Measure of the System Dependence...”、“Call for Rigor...”）。人们对现有方法持有一种健康的怀疑态度，并致力于实现更鲁棒、可靠和有意义的进展评估。对创建超越简单学术任务、具有多样性和挑战性的数据集（如 <b>1</b> 中的“HateDay”、“AfriMed-QA”）的强调，反映了该领域的成熟，即仅仅在现有基准上追求更高分数已被认为是不够的。人们要求评估能够更好地反映实际效用和更深层次的理解。因此，数据集和评估协议的质量和性质本身就是活跃的研究领域。这种自我反思和批判的立场对于确保真正的科学进步和避免虚幻的成果至关重要。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b24580dcaa36d6be60311408" data-id="208d0968b24580dcaa36d6be60311408"><span><div id="208d0968b24580dcaa36d6be60311408" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b24580dcaa36d6be60311408" title="5. 跨学科趋势分析与未来展望"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>5. 跨学科趋势分析与未来展望</b></span></span></h3><div class="notion-text notion-block-208d0968b24580c782d3d631f5ad85de">综合上述对ACL 2025主要研究方向的分析，可以观察到一些显著的跨学科趋势和未来发展方向。这些趋势往往是多个研究领域交叉融合的结果，预示着NLP技术未来的演进路径。</div><div class="notion-text notion-block-208d0968b24580fd8836ccf6fb247a8c"><b>新兴模式:</b></div><ul class="notion-list notion-list-disc notion-block-208d0968b245807597b4dd806d37d801"><li><b>AI的智能体化 (Agentification of AI)</b>：一个非常清晰的趋势是发展由LLM驱动的智能体，它们能够执行复杂任务、使用工具并与环境交互。论文如《MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments》、《CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System》、《LegalAgentBench: Evaluating LLM Agents in Legal Domain》和《SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social Media》均体现了这一方向 。</li><ul class="notion-list notion-list-disc notion-block-208d0968b245807597b4dd806d37d801"><div class="notion-text notion-block-208d0968b245800da06cfe673cba3518"><b>1</b></div></ul></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580bebb76f912631e5d1b"><li><b>个性化 (Personalization)</b>：根据个体用户或特定情境定制LLM的行为和响应，正成为一个重要的研究领域。例如，《Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas》和《PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling》等工作致力于实现更个性化的NLP系统 。</li><ul class="notion-list notion-list-disc notion-block-208d0968b24580bebb76f912631e5d1b"><div class="notion-text notion-block-208d0968b24580289beae13069958672"><b>1</b></div></ul></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245808fa3c7f4f3b0f3223d"><li><b>人机协作 (Human-AI Collaboration)</b>：设计使人类和AI能够协同工作的系统，是另一个日益受到关注的领域。相关研究包括综述性工作《How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond》以及具体的框架设计《Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration》。</li><ul class="notion-list notion-list-disc notion-block-208d0968b245808fa3c7f4f3b0f3223d"><div class="notion-text notion-block-208d0968b2458073a7b0f703f3838c9f"><b>1</b></div></ul></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245800d9f6ed63a1c293af9"><li><b>“元”层面研究 (The &quot;Meta&quot; Layer)</b>：对LLM本身（可解释性、探针分析、理论性质）以及研究过程（评估方法、数据质量）的关注显著增加，显示出领域对自身基础和方法论的深刻反思。</li></ul><div class="notion-text notion-block-208d0968b245807bb320c47baf0ab119"><b>潜在未来轨迹:</b></div><ul class="notion-list notion-list-disc notion-block-208d0968b245802bae9ef9e7b1c22ebc"><li>持续推动更鲁棒和可泛化的推理能力。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580df9da4d3b57b95741e"><li>符号推理与神经方法的更紧密集成。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b245808b8acee425755a2e79"><li>更复杂的多模态锚定和交互。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580859482d2c5d0fbedaa"><li>在极低资源NLP方面取得突破。</li></ul><ul class="notion-list notion-list-disc notion-block-208d0968b24580be853adf73056f7e5a"><li>开发可验证且本质安全的AI系统。</li></ul><div class="notion-text notion-block-208d0968b245801896b4cbce0cb33e93">更广泛的启示：NLP作为普适性赋能技术</div><div class="notion-text notion-block-208d0968b24580428e42eff7a13b450e">ACL 2025所展现的研究趋势表明，NLP正从一个专门的学科学科转变为一项核心的赋能技术，其应用范围横跨科学、工业和社会的各个方面。从医疗（如 1 中的“AfriMed-QA”）、法律（如 1 中的“LegalAgentBench”）、金融（如 1 中的“INVESTORBENCH”）、电商（如 1 中的“EcomScriptBench”），到社会科学（如 1 中的“When People are Floods...”）和代码工程（如 1 中的“WarriorCoder”），NLP的应用场景日益广泛。同时，LLM正被开发为能够在数字甚至物理环境中行动的智能体（如 1 中的“CompileAgent”、“AndroidLab”）。这种应用的广度以及向智能体化的发展表明，NLP不再仅仅是处理文本，而是关于创建能够基于语言和多模态输入进行理解、推理和行动的智能系统。随着NLP工具变得越来越强大和普及，其社会影响（无论是积极的还是消极的）都将随之增长。这进一步凸显了对伦理、安全和可控性研究的极端重要性。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-208d0968b2458031a158c6241dbf65e5" data-id="208d0968b2458031a158c6241dbf65e5"><span><div id="208d0968b2458031a158c6241dbf65e5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#208d0968b2458031a158c6241dbf65e5" title="6. 结论：反思ACL 2025对NLP的贡献"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><b>6. 结论：反思ACL 2025对NLP的贡献</b></span></span></h3><div class="notion-text notion-block-208d0968b24580caab4dc6d98d4de969">ACL 2025的论文清晰地勾勒出自然语言处理领域当前的研究热点和未来走向。大型语言模型（LLMs）无疑是推动本轮NLP浪潮的核心引擎，其影响无处不在。整个研究社群正致力于充分发挥LLM的巨大潜力，同时积极应对其带来的各种挑战，包括提升其核心能力、确保其可靠性、降低其应用门槛、以及规范其伦理边界。</div><div class="notion-text notion-block-208d0968b245800d809fde143f136404">会议论文集展示了NLP领域在多个前沿方向上的积极探索：从更深层次的推理和可控生成，到更高效、更安全的模型架构；从更全面的评测基准和方法论，到更广泛的多语言覆盖和文化适应；从更逼真的多模态交互，到更智能的信息获取和知识融合。特别值得注意的是，面向特定领域的应用研究以及智能体系统的兴起，预示着NLP技术正加速从实验室走向真实世界，赋能千行百业。</div><div class="notion-text notion-block-208d0968b245800089f7eb048ba38795">然而，伴随着技术的飞速发展，对伦理、公平、透明和可控性的关注也达到了前所未有的高度。ACL 2025的研究成果反映出社群在追求技术创新的同时，也在努力构建负责任的AI生态系统。</div><div class="notion-text notion-block-208d0968b245806cad27f98ac4a374fd">总体而言，ACL 2025不仅是NLP最新研究成果的展示平台，更是领域内思想碰撞、方向引领的重要场域。它揭示了一个充满活力、快速演进，并勇于面对核心挑战的NLP研究社群。在快速创新与严谨、负责任的开发之间取得平衡，将是NLP领域未来持续健康发展的关键。</div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[网络文本情感计算]]></title>
            <link>http://preview.tangly1024.com/article/206d0968-b245-80e9-985f-c5552cd13350</link>
            <guid>http://preview.tangly1024.com/article/206d0968-b245-80e9-985f-c5552cd13350</guid>
            <pubDate>Mon, 02 Jun 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[网络文本情感计算推荐论文。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-206d0968b24580e9985fc5552cd13350"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><div class="notion-text notion-block-206d0968b245805babeddc2d9e80e586">基础版</div><ol start="1" class="notion-list notion-list-numbered notion-block-206d0968b2458012bbbcf0bdb9e9d9b9"><li>Mining Opinion Features in Customer Reviews</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-206d0968b2458051a6f6e2def2603d9f"><li>Mining and summarizing customer reviews</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-206d0968b24580808f31ee6cfcdcbc0b"><li><a class="notion-link" href="/206d0968b2458050b4c8c7a3395688b5">A Holistic Lexicon-Based Approach to Opinion Mining</a></li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-206d0968b2458011bc53e0e4208c5e62"><li>Expanding Domain Sentiment Lexicon through Double Propagation</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-206d0968b24580dc8718e96d1f9f2723"><li>Opinion Word Expansion and Target Extraction through Double Propagation</li></ol><ol start="6" class="notion-list notion-list-numbered notion-block-206d0968b245800684cfd1db21e26a75"><li>Thumbs up? Sentiment Classification using Machine Learning Techniques</li></ol><ol start="7" class="notion-list notion-list-numbered notion-block-206d0968b24580029485e1ca912bb85e"><li>A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts</li></ol><ol start="8" class="notion-list notion-list-numbered notion-block-206d0968b2458009a298cfd240174ac5"><li>Recognizing contextual polarity in phrase-level sentiment analysis</li></ol><ol start="9" class="notion-list notion-list-numbered notion-block-206d0968b24580f5a557e7dadb2818a1"><li>Twitter sentiment analysis: The good the bad and the omg!</li></ol><ol start="10" class="notion-list notion-list-numbered notion-block-206d0968b2458094aa59d6f91906015e"><li>Twitter sentiment classification using distant supervision</li></ol><ol start="11" class="notion-list notion-list-numbered notion-block-206d0968b2458062a8b3efda9ad78032"><li>Learning extraction patterns for subjective expressions</li></ol><div class="notion-text notion-block-206d0968b2458088812df93b428b1df4">高级版</div><div class="notion-text notion-block-206d0968b24580d391e2d71912ac0630">论文列表：</div><div class="notion-text notion-block-206d0968b245800f95f7d66d921e9c6b">（1）Convolutional Neural Networks for Sentence Classification.</div><div class="notion-text notion-block-206d0968b245809eb68ad4051dd2c48d">（2）Hierarchical Attention Networks for Document Classification.</div><div class="notion-text notion-block-206d0968b24580f1881afafe510c1800">（3）《Efficient Estimation of Word Representations in Vector Space》(original word2vec paper)</div><div class="notion-text notion-block-206d0968b245801c88ccc1d2855a3d4e">（4）《GloVe: Global Vectors for Word Representation》（original GloVe paper）</div><div class="notion-text notion-block-206d0968b245806bbca6f0577aec4fce">（5）《Sequence Modeling: Recurrent and Recursive Neural Nets》(Sections 10.1 and 10.2)</div><div class="notion-text notion-block-206d0968b2458022a0cfc1ba828cd856">（6）《Learning long-term dependencies with gradient descent is difficult》(one of the original vanishing gradient papers)</div><div class="notion-text notion-block-206d0968b2458067844bc95b684dd882">（7）《Neural Machine Translation by Jointly Learning to Align and Translate》(original seq2seq+attention paper)</div><div class="notion-text notion-block-206d0968b24580289747d83258cdb5ce">（8）《Practical Methodology》(Deep Learning book chapter)</div><div class="notion-text notion-block-206d0968b2458065ba96f0f24d3aed0f">（9）《Attention Is All You Need》</div><div class="notion-text notion-block-206d0968b2458085b546c5131f1aab3e">（10）《BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding》</div><div class="notion-text notion-block-206d0968b2458008a9a8f113ff7002f4">（11）《Reading Wikipedia to Answer Open-Domain Questions》</div><div class="notion-text notion-block-206d0968b24580c9a4d2d7f629156536">（12）《Get To The Point: Summarization with Pointer-Generator Networks》</div><div class="notion-text notion-block-206d0968b24580bfb439c81a31e735a1">（13）《End-to-end Neural Coreference Resolution.》</div><div class="notion-text notion-block-206d0968b2458078a4a5dc8c7b781881">（14）《Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer》</div><div class="notion-text notion-block-206d0968b2458097b86efc28bb6cb1d1">（15）《Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model》</div><div class="notion-text notion-block-206d0968b245802191b0dc85477e63c0">实体识别：实体识别是将文本中的实体与知识图谱中的实体进行匹配的过程。常用的实体识别算法有命名实体识别(Named Entity Recognition，NER)和实体链接(Entity Linking，EL)。</div><div class="notion-text notion-block-206d0968b245803c8c65e26b63273d3b">关系抽取：关系抽取是从文本中抽取实体之间的关系的过程。常用的关系抽取算法有规则基于的方法(Rule-based Method)和机器学习基于的方法(Machine Learning Based Method)。</div><div class="notion-text notion-block-206d0968b24580fbbafbc709c843b0fe">情感分析：情感分析是将文本中的情感倾向与实体进行关联的过程。常用的情感分析算法有基于词汇量的方法(Lexicon-based Method)和深度学习基于的方法(Deep Learning Based Method)。</div><div class="notion-text notion-block-206d0968b245804c9303fb0f17a350b0">原文链接：https://blog.csdn.net/universsky2015/article/details/135787422</div></main></div>]]></content:encoded>
        </item>
    </channel>
</rss>