<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>HMM on 111qqz's blog</title><link>https://111qqz.com/en/tags/hmm/</link><description>Recent content in HMM on 111qqz's blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>hust.111qqz@gmail.com (111qqz)</managingEditor><webMaster>hust.111qqz@gmail.com (111qqz)</webMaster><copyright>© 2011-2026 111qqz</copyright><lastBuildDate>Sat, 10 Oct 2026 14:30:00 +0800</lastBuildDate><atom:link href="https://111qqz.com/en/tags/hmm/index.xml" rel="self" type="application/rss+xml"/><item><title>从马尔可夫性质出发：DP 无后效性、HMM 与 VAE 的知识联想</title><link>https://111qqz.com/en/post/%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/2026-10-10-markov-property-knowledge-associations/</link><pubDate>Sat, 10 Oct 2026 14:30:00 +0800</pubDate><author>hust.111qqz@gmail.com (111qqz)</author><guid>https://111qqz.com/en/post/%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/2026-10-10-markov-property-knowledge-associations/</guid><description>&lt;h2 class="relative group"&gt;起因
 &lt;div id="起因" class="anchor"&gt;&lt;/div&gt;
 
 &lt;span
 class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none"&gt;
 &lt;a class="text-primary-300 dark:text-neutral-700 !no-underline" href="#%e8%b5%b7%e5%9b%a0" aria-label="Anchor"&gt;#&lt;/a&gt;
 &lt;/span&gt;
 
&lt;/h2&gt;
&lt;p&gt;继续跟着 MIT 6.S191 Lecture 5 学强化学习。&lt;a href="https://111qqz.com/en/post/%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/2026-10-09-rl-return-to-q-function/" &gt;上一篇&lt;/a&gt;理解了 Return、State Value Function \(V^\pi(s)\) 和 Action Value Function \(Q^\pi(s,a)\)。&lt;/p&gt;
&lt;p&gt;接下来要学的是 Bellman Equation。&lt;/p&gt;
&lt;p&gt;我之前隐约知道它能把 \(V^\pi(s)\) 写成递归形式——当前状态的长期价值，等于下一步的期望 Reward 加上折扣后下一状态的期望长期价值。但刚准备细看推导时，发现它有一个前提条件：&lt;/p&gt;</description></item></channel></rss>