Intel 软件工程师,主要负责开源大语言模型推理框架在 Intel 平台上的适配和性能调优工作。
大型语言模型 (LLM) 凭借其卓越的能力吸引了来自工业界和学术界的广泛关注。然而,一个主要障碍在于,现存的大部分 LLM 研究和开发工作都依赖于 GPU,但价格昂贵且供應有限并非人人得享用。为了弥合这一差距,我们正在积极致力于让一些流行的 LLM 框架能够高效地运行在 CPU 上。这将使这些强大的工具能够被更广泛的使用,最终加快大型语言模型领域的进步。在这次演讲中,我们以 vLLM 为例,分享我们在 Intel Xeon 上实现最大化应用程序性能的最佳实践。
The world of Large Language Models (LLMs) has sparked significant interest in both industry and academia due to their remarkable capabilities. However, a major hurdle exists particularly in China: most current LLM research and development relies heavily on GPUs, which are expensive and not always accessible. To bridge this gap, we are actively working on enabling some popular LLM frameworks to run efficiently on CPUs. This will make these powerful tools more accessible to a wider range of users, ultimately accelerating advancements in the field of LLMs. Using vLLM as an example, we share our best practices to maximize inference performance on Intel Xeon.
演讲提纲:
1. vLLM 是業界領先的开源服务框架
2. Intel® 与 vLLM 社区合作集成 vLLM 在Xeon® 服务器的開發和性能优化
3. RFC: https://github.com/vllm-project/vllm/issues/3654
4. 文档: https://docs.vllm.ai/en/latest/getting_started/cpu-installation.html
5. Intel® Xeon® 服务器是 LLM 推理一个不错的选择
6. Q & A
听众收益:



微信咨询

电话咨询
微信联系我们

