Resiliency at Scale: Managing Google's TPUv4 Machine Learning Supercomputer.

Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles,Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, Henri Bahini

Symposium on Networked Systems Design and Implementation(2024)

引用 0|浏览7
暂无评分
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要