
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Kubernetes Administration – Praxiswissen &amp; Best Practices</title>
      <link>https://www.kubernetes-administration.de/blog</link>
      <description>Praxisnahe Anleitungen, Best Practices und News rund um Kubernetes Administration, Betrieb und Sicherheit – verständlich für Engineers und Teams.</description>
      <language>de-de</language>
      <managingEditor>kontakt@pexon-consulting.de (Phillip Pham)</managingEditor>
      <webMaster>kontakt@pexon-consulting.de (Phillip Pham)</webMaster>
      <lastBuildDate>Sun, 30 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.kubernetes-administration.de/tags/inference/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.kubernetes-administration.de/blog/agentic-ai-auf-kubernetes-praxis</guid>
    <title>Agentic AI auf Kubernetes: Was in der Praxis zählt</title>
    <link>https://www.kubernetes-administration.de/blog/agentic-ai-auf-kubernetes-praxis</link>
    <description>Agentic AI auf Kubernetes: Warum K8s das Substrat bleibt, welche Schichten (Inference bis Agents) zählen und wo Sandbox, Scheduling und Kosten knifflig werden.</description>
    <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
    <author>kontakt@pexon-consulting.de (Phillip Pham)</author>
    <category>kubernetes</category><category>ai-agents</category><category>inference</category><category>gpu</category><category>platform-engineering</category><category>aks</category>
  </item>

  <item>
    <guid>https://www.kubernetes-administration.de/blog/kubernetes-ai-at-scale-dra-inference</guid>
    <title>Kubernetes AI at Scale: DRA, LLMD und Inference</title>
    <link>https://www.kubernetes-administration.de/blog/kubernetes-ai-at-scale-dra-inference</link>
    <description>Kubernetes wird Accelerator Native: DRA, LLMD, Disaggregated Serving und Inference Gateway für produktive GenAI-Workloads — was Plattform-Teams jetzt brauchen.</description>
    <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
    <author>kontakt@pexon-consulting.de (Phillip Pham)</author>
    <category>kubernetes</category><category>ai</category><category>dra</category><category>inference</category><category>gpu</category><category>llm</category><category>platform-engineering</category>
  </item>

  <item>
    <guid>https://www.kubernetes-administration.de/blog/llm-d-distributed-inference-kubernetes</guid>
    <title>LLM-D: Verteilte Inference auf Kubernetes</title>
    <link>https://www.kubernetes-administration.de/blog/llm-d-distributed-inference-kubernetes</link>
    <description>LLM-D verteilt Inference über Kubernetes: Prefill/Decode trennen, KV-Cache nutzen, intelligent routen — niedrigere Latenz und bessere GPU-Kosten.</description>
    <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
    <author>kontakt@pexon-consulting.de (Phillip Pham)</author>
    <category>kubernetes</category><category>llm</category><category>inference</category><category>vllm</category><category>ai</category><category>gpu</category>
  </item>

    </channel>
  </rss>
