Polars 2.0 Release Candidate Makes the Streaming Engine the Default, Breaking CSV Reads and Type Coercion
Polars 2.0-rc.1 flips LazyFrame queries to the streaming engine by default, a change the project says should be roughly 5x faster while removing melt(), join_nulls, and other older APIs.
Editor's Note ·
- Correction:
- The article states the migration guide marks LazyFrame.profile() as removed because it is "incompatible with streaming engine." That exact phrase does not appear in the migration guide. The guide's actual wording is: "It was designed for the in-memory engine; the concurrent nature of the streaming engine, now the default, would make its per-node timings misleading."
Overview
Polars, the Rust-built “Extremely fast Query Engine for DataFrames, written in Rust”, shipped the first release candidate for Polars 2.0 on September 2, 2026, according to Polars. The change at the center of the major-version bump is narrow but consequential: calling collect on a LazyFrame now defaults to Polars’ streaming engine instead of its in-memory engine, a switch the announcement says should “be easily 5x faster” in aggregate, according to Polars.
What We Know
- The core breaking change is that “calling
collecton aLazyFramewill now default to the streaming engine, leading to massive memory and performance improvements on most queries for users,” per Polars. - The major-version bump is required specifically because the streaming engine “doesn’t guarantee row-order by default for certain operations (
join,group_by,unpivot, etc.),” according to Polars. - The release notes for the py-2.0.0-rc.1 tag on GitHub list “Set the default engine for SQL to the streaming engine” as the sole entry under breaking changes, tracked in pull request #28973.
- The official migration guide documents a wider set of removed or renamed APIs alongside the engine switch:
melt()is replaced byunpivot(), thejoin_nullsparameter is replaced bynulls_equal, andpl.read_csv()now internally dispatches topl.scan_csv(...).collect(). - The migration guide also marks
LazyFrame.profile()as removed because it is “incompatible with streaming engine,” and it dropsDataFrame.__dataframe__(), ending support for the DataFrame Interchange Protocol. - Concatenation becomes stricter under the new defaults:
pl.concat(..., how="horizontal")now requires equal-height frames by default rather than silently padding shorter ones with nulls, per the migration guide. - Beyond the engine-default change, the GitHub release notes catalog a set of SQL- and Iceberg-focused additions, including reused native metadata for Iceberg sinks, predicate pushdown for SQL
EXISTSsubqueries, join reordering, and support for partitioned Iceberg sinks. - Developers can try the release candidate with
pip install polars==2.0rc1, according to Polars, and the migration guide describes a configuration option to opt back into the pre-2.0 in-memory execution model for queries that need it.
What We Don’t Know
- The announcement and release notes reviewed do not state a target date for the final, non-release-candidate Polars 2.0 build.
- The materials do not detail how teams that relied on the now-removed
LazyFrame.profile()method for performance debugging are expected to adapt, beyond the migration guide’s note that the method is incompatible with the streaming engine.
Analysis
Polars’ own framing casts 2.0 as less of a feature release and more of a defaults cleanup, but the practical effect for existing pipelines is still a major version bump: the streaming engine’s lack of guaranteed row order for join, group_by, and unpivot means any code that implicitly depended on output ordering from those operations can silently start returning rows in a different sequence once a project upgrades. That the row-order tradeoff alone was judged enough to require a major version — rather than being folded into a minor release — underscores how much weight the project places on giving users an explicit signal before changing execution semantics that downstream code may quietly rely on.