Introducing English as the New Programming Language for Apache Spark

所属专题:从 BI 到 BI+AI,新计算范式下的大数据平台

嘉宾 : 李潇 | Databricks 工程总监、Apache Spark PMC

会议室 : 首府 5

讲师介绍

专题演讲嘉宾:李潇

Databricks 工程总监、Apache Spark PMC

Xiao Li is an engineering director, Apache Spark committer and PMC member at Databricks. He is leading and managing seven teams for development of Apache Spark, Databricks Runtime and DB SQL. His main interests are on data lakehouse, data replication and data integration. Previously, he was an IBM master inventor and an expert on asynchronous database replication and consistency verification. He received his Ph.D. from University of Florida in 2011.

Xiao Li 是 Databricks 的工程总监、Apache Spark Committer 和 PMC 成员。他领导和管理七个团队,负责开发 Apache Spark、Databricks Runtime 和 DB SQL。他的主要兴趣是数据湖仓、数据复制和数据集成。此前,他是 IBM Master Inventor 荣誉的获得者,也是数据库异步复制和一致性验证方面的专家。他于 2011 年在佛罗里达大学获得博士学位。

议题介绍

演讲:Introducing English as the New Programming Language for Apache Spark

In this talk, we will introduce the English SDK for Apache Spark, a pioneering tool developed to enhance the accessibility and usability of Apache Spark through the innovative use of Generative AI. The goal is to transform the conventional programming paradigm, shifting from AI as a co-pilot to a chauffeur, thereby making Spark more approachable and user-friendly.

Addressing the limitations in AI-assisted code development, like GitHub Copilot's occasional struggle with context especially with Spark tables and DataFrames, the English SDK introduces the concept of English as a programming language. With the aid of Generative AI, English instructions are compiled into PySpark and SQL code, reducing the necessity for users to comprehend complex code structures. An example is its ability to execute DataFrame transformations via a simple English instruction.

This novel approach extends to key features such as data ingestion, DataFrame operations, user-defined functions (UDFs), and caching. Notably, the SDK can search, select, and integrate web data into Spark in a single step; simplify DataFrame operations through intuitive English descriptions; streamline the creation of UDFs with AI-assisted code completion; and incorporate caching to improve execution speed, reproducibility, and cost efficiency.

Overall, the English SDK, built on the community’s extensive contributions to Spark, represents a transformative leap towards making data analytics more accessible, advancing our objective to broaden the reach of Apache Spark.

交通指南

北京·富力万丽酒店

Renaissance Beijing Hotel
地址:北京市朝阳区东三环中路61号
  • 微信咨询

  • 电话咨询

    联系电话:+86 18514549229

领取往期热门演讲视频

领取往期热门演讲视频二维码
如您在购票过程中遇到问题,请扫码咨询票务小姐姐