Xinyu Liu
Senior Software Engineer, Alibaba Storage Services Team; Paimon C++ Lead
Xinyu Liu is a Senior Software Engineer on the Storage Services Team at Alibaba, an Apache Paimon Committer, and the Maintainer of Paimon-cpp. He focuses on the core development and open-source maintenance of Paimon-cpp and has successfully driven its adoption in Alibaba’s internal data lake. Previously, he worked on the development of the time-series database Khronos, providing data storage and query capabilities for Alibaba Group’s monitoring systems. His research interests include storage engines, high-performance data formats, and system-level performance optimization. He is committed to working with the Apache community to build a native data lakehouse ecosystem.
Topic
Paimon-cpp: Apache Paimon C++ Implementation and Production Practices for Native Query Engines
Apache Paimon has become one of the most actively developing data lakehouse formats in the Apache ecosystem. However, current mainstream implementations mainly rely on the JVM, presenting new challenges for the growing number of native query engines that pursue low latency, high performance, and ease of embedding. This presentation will introduce the Paimon-cpp project, which we led the development of, and explain how to build a high-performance C++ implementation compatible with Apache Paimon behavior from scratch. It supports Append Table, Primary Key Table, MOR, Deletion Vector, Data Evolution, multiple types of indexes, and real-time data query capabilities. We will focus on the core design based on the Arrow columnar memory format, as well as how to improve read/write and Compaction performance through shallow-copy data exchange, prefetching optimization, asynchronous parallel row-column conversion, Blob I/O optimization, and Map shared columnization design for time-series data. We will also introduce how its modular plugin-based architecture can adapt to different file systems, file formats, thread pools, and metrics components. Finally, based on practical engineering experience, we will share key practices and lessons learned from integrating Paimon-cpp with native engines, production optimization, and open-source contributions. This presentation will start with the development background and design goals of Paimon-cpp, introducing how it maintains compatibility with Apache Paimon behavior and supports Append Table, Primary Key Table, MOR, Deletion Vector, Data Evolution, multiple types of indexes, and real-time data queries. It will further share the Arrow-based columnar memory design, as well as key implementations including shallow-copy data exchange, prefetching optimization, asynchronous parallel row-column conversion, Blob I/O optimization, and Map shared columnization for time-series data. Finally, based on practical implementation experience, it will introduce its modular architecture design, production optimization, and native engine integration practices. Attendees will gain an understanding of the core challenges and implementation approaches involved when integrating lakehouse formats into native query engines, systematically understand the format compatibility, performance optimization, and architectural design approaches of the Paimon C++ implementation, and gain reusable practical experience in real-time querying, primary key table processing, time-series data modeling, and production deployment. This will provide useful references for building self-developed lakehouse engines, storage systems, or data platforms.