<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Self-Attention on 111qqz的小窝</title><link>https://111qqz.com/tags/self-attention/</link><description>Recent content in Self-Attention on 111qqz的小窝</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>hust.111qqz@gmail.com (111qqz)</managingEditor><webMaster>hust.111qqz@gmail.com (111qqz)</webMaster><copyright>© 2011-2026 111qqz</copyright><lastBuildDate>Sun, 13 Sep 2026 20:00:00 +0800</lastBuildDate><atom:link href="https://111qqz.com/tags/self-attention/index.xml" rel="self" type="application/rss+xml"/><item><title>去掉 RNN 之后，顺序去哪了？</title><link>https://111qqz.com/2026/09/where-sequence-order-went/</link><pubDate>Sun, 13 Sep 2026 20:00:00 +0800</pubDate><author>hust.111qqz@gmail.com (111qqz)</author><guid>https://111qqz.com/2026/09/where-sequence-order-went/</guid><description>&lt;p&gt;上一篇&lt;a href="https://111qqz.com/2026/09/rnn-to-self-attention/" &gt;《从 RNN 到 Self-Attention》&lt;/a&gt;结尾记账时留了一条：去掉 recurrence 之后，顺序不再天然编码在计算过程里，摊平的 attention 层需要额外把顺序信息补回去，「具体怎么补，下一篇再说」。&lt;/p&gt;</description></item><item><title>从 RNN 到 Self-Attention：信息为什么一定要沿着 State 一步步传递？</title><link>https://111qqz.com/2026/09/rnn-to-self-attention/</link><pubDate>Sun, 13 Sep 2026 18:00:00 +0800</pubDate><author>hust.111qqz@gmail.com (111qqz)</author><guid>https://111qqz.com/2026/09/rnn-to-self-attention/</guid><description>&lt;p&gt;上一篇&lt;a href="https://111qqz.com/2026/09/seq2seq-to-attention/" &gt;《从 RNN / LSTM / GRU 到早期 Attention》&lt;/a&gt;结尾我留了一句话：早期的 attention 并不是为了取代 RNN 而出现的，它只是夹在 encoder–decoder 中间解决 fixed-context bottleneck；但“根据当前需求去读取一组 representation”这种机制，并不一定要依附在 RNN 上。这篇接着讲那个“另一个故事”。&lt;/p&gt;</description></item></channel></rss>