Best Practices of running vLLM on Xeon

所属专题:性能优化

嘉宾 : (1) Tony Wu | Intel 机器学习性能高级性能架构师(2) 李江 | Intel软件工程师

会议室 : 国盛厅

讲师介绍

专题演讲嘉宾:Tony Wu

Intel 机器学习性能高级性能架构师

专题演讲嘉宾:李江

Intel软件工程师

Intel 软件工程师,主要负责开源大语言模型推理框架在 Intel 平台上的适配和性能调优工作。

议题介绍

地点:国盛厅
所属专题:性能优化

演讲:Best Practices of running vLLM on Xeon

大型语言模型 (LLM) 凭借其卓越的能力吸引了来自工业界和学术界的广泛关注。然而,一个主要障碍在于,现存的大部分 LLM 研究和开发工作都依赖于 GPU,但价格昂贵且供應有限并非人人得享用。为了弥合这一差距,我们正在积极致力于让一些流行的 LLM 框架能够高效地运行在 CPU 上。这将使这些强大的工具能够被更广泛的使用,最终加快大型语言模型领域的进步。在这次演讲中,我们以 vLLM 为例,分享我们在 Intel Xeon 上实现最大化应用程序性能的最佳实践。

The world of Large Language Models (LLMs) has sparked significant interest in both industry and academia due to their remarkable capabilities. However, a major hurdle exists particularly in China: most current LLM research and development relies heavily on GPUs, which are expensive and not always accessible. To bridge this gap, we are actively working on enabling some popular LLM frameworks to run efficiently on CPUs. This will make these powerful tools more accessible to a wider range of users, ultimately accelerating advancements in the field of LLMs. Using vLLM as an example, we share our best practices to maximize inference performance on Intel Xeon.

演讲提纲:

1. vLLM 是業界領先的开源服务框架

2. Intel® 与 vLLM 社区合作集成 vLLM 在Xeon® 服务器的開發和性能优化

3. RFC: https://github.com/vllm-project/vllm/issues/3654

4. 文档: https://docs.vllm.ai/en/latest/getting_started/cpu-installation.html

5. Intel® Xeon® 服务器是 LLM 推理一个不错的选择

6. Q & A

听众收益:

  •  Help audience better understand pros and cons for running AI vLLM on Xeon
  •  Help audience develop knowledge of LLM inference workload behavior on Xeon
  •  Help audience tune and optimize LLM serving performance on Xeon
  •  帮助观众更好地理解 在 Xeon 平台上運行vLLM的利弊
  •  帮助观众拓展 vLLM 工作负载行为的相关知识
  •  帮助观众开发和优化大語言模型性能,在 Intel Xeon 上获得更好的性能

交通指南

北京国测国际会议会展中心

GUOCE International Convention And Exhibition Center
地址:北京市顺义区临空经济核心区汇海南路6号院20号楼
  • 微信咨询

  • 电话咨询

    联系电话:+86 17310043226

微信联系我们

如您在购票过程中遇到问题,请扫码咨询票务小助手