SemiAnalysis 测试 Vultr GPU 集群暴露多项问题
SemiAnalysis 在 Vultr GPU 集群上测试时发现多项问题:镜像含 CUDA、Docker、runc 等多个存在 CVE 的库,登录 pod 缺少 `/shared` 挂载和 Pyxis 配置,K8s 默认 `vultr-block-storage` StorageClass 与裸金属集群不兼容导致 pod 卡在 `ContainerCreating`。
The images that Vultr built for our testing had many libraries with applicable CVEs, including CUDA, DCGM, Docker, runc, and ConnectX firmware. Our Slinky login pod did not mount `/shared` and lacked the Pyxis plugstack config, so we had to orchestrate from a worker pod until the Vultr support team reprovisioned the login pod mid-campaign, fixing both issues. There was another basic problem with the storage visible from the K8s layer: the default storage configuration was incompatible with our bare-metal cluster. The default StorageClass, `vultr-block-storage`, provisioned and bound without complaint, then the pod sat in `ContainerCreating` forever. However, the Lustre tier, which came attached to the worker nodes, was excellent, achieving 91.7 GB/s aggregate read across 32 clients. Grafana was configured but only partially correct, as NVLink read 0 due to unconfigured `DCGM_FI_PROF_NVLINK_*` fields. (2/3)
来源:SemiAnalysis_ · x.com