<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Deep-Learning on 111qqz的小窝</title><link>https://111qqz.com/tags/deep-learning/</link><description>Recent content in Deep-Learning on 111qqz的小窝</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>hust.111qqz@gmail.com (111qqz)</managingEditor><webMaster>hust.111qqz@gmail.com (111qqz)</webMaster><copyright>© 2011-2026 111qqz</copyright><lastBuildDate>Sun, 11 Oct 2026 15:45:00 +0800</lastBuildDate><atom:link href="https://111qqz.com/tags/deep-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>从 Value Iteration 到 Q-Learning：没有真实 Q 值，模型如何学会决策？</title><link>https://111qqz.com/2026/10/value-iteration-to-q-learning-dqn/</link><pubDate>Sun, 11 Oct 2026 15:45:00 +0800</pubDate><author>hust.111qqz@gmail.com (111qqz)</author><guid>https://111qqz.com/2026/10/value-iteration-to-q-learning-dqn/</guid><description>&lt;p&gt;&lt;a href="https://111qqz.com/2026/10/policy-improvement-to-value-iteration/" &gt;上一篇&lt;/a&gt;结尾留了一个问题：如果不知道环境的转移概率和 Reward 函数，应该怎么利用 Bellman Equation 学习策略？&lt;/p&gt;
&lt;p&gt;那篇从 Policy Iteration 走到 Value Iteration，用 A/B 两状态的例子跑完了整个过程。最后得到了最优策略，A 选择前往 B，B 选择 \(b_1\)，\(V^*(A) = 2.7\)。&lt;/p&gt;</description></item></channel></rss>