实验 11: 使用 HAMi DRA 共享 GPU 运行 KServe 推理服务部署 KServe Standard vLLM 服务,并通过 HAMi 原生 DRA Claim 让两个 Predictor 副本共享一张 NVIDIA GPU。