<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Performance-Engineering on 111qqz的小窝</title><link>https://111qqz.com/tags/performance-engineering/</link><description>Recent content in Performance-Engineering on 111qqz的小窝</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>hust.111qqz@gmail.com (111qqz)</managingEditor><webMaster>hust.111qqz@gmail.com (111qqz)</webMaster><copyright>© 2011-2026 111qqz</copyright><lastBuildDate>Tue, 15 Sep 2026 22:53:41 +0800</lastBuildDate><atom:link href="https://111qqz.com/tags/performance-engineering/index.xml" rel="self" type="application/rss+xml"/><item><title>模型压缩之后，为什么推理反而变慢了：一次 CPU Serving 的性能优化实践</title><link>https://111qqz.com/2026/09/cpu-serving-decompression-tax/</link><pubDate>Tue, 15 Sep 2026 22:53:41 +0800</pubDate><author>hust.111qqz@gmail.com (111qqz)</author><guid>https://111qqz.com/2026/09/cpu-serving-decompression-tax/</guid><description>&lt;p&gt;之前我负责过一个推荐系统深度学习 Serving 框架的功能开发与性能优化。&lt;/p&gt;
&lt;p&gt;我们的系统主要跑在 CPU 上，以 TensorFlow 为主力 inference engine。随着模型迭代，尤其是 Embedding 层越来越大，模型导出、传输和上线变得非常慢，Serving 节点的内存成本也眼看着往上飙。&lt;/p&gt;</description></item></channel></rss>