
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Vegetog&#39;s Blog</title>
      <link>https://blog.vegetog.com/blog</link>
      <description>daily routine and something</description>
      <language>zh-CN</language>
      <managingEditor>undefined (Vegetog)</managingEditor>
      <webMaster>undefined (Vegetog)</webMaster>
      <lastBuildDate>Thu, 10 Sep 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://blog.vegetog.com/tags/transformer/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://blog.vegetog.com/blog/ai-infra/prefill-parallelism-and-bottlenecks</guid>
    <title>关于 Prefill 一些模糊点的理解</title>
    <link>https://blog.vegetog.com/blog/ai-infra/prefill-parallelism-and-bottlenecks</link>
    <description>从 token 与 attention head 的并行、矩阵化计算中的权重复用，到算术强度和耗时下界，厘清 prefill 与 decode 的性能差异。</description>
    <pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate>
    <author>undefined (Vegetog)</author>
    <category>AI Infra</category><category>LLM</category><category>Transformer</category>
  </item>

    </channel>
  </rss>
