library(sparklyr)
sc <- spark_connect(master = "local", method = "sail")
#> Retrieving version from PyPi.org
#> ✔ PyPi specs: 'pyspark-client' version 4.2.0, requires Python >=3.10 [50ms]
#>
#> ℹ
#> ✔ Python environment: 'Managed `uv` environment' [943ms]
#>
#> [2026-10-05T20:20:48Z INFO sail_python::spark::server] Starting the Spark Connect server on 127.0.0.1:57003...
#> [2026-10-05T20:20:48Z INFO sail_session::session_manager::actor::handler] creating session 1fa24ec4-596b-4746-826e-7013e9e2fa3bSail
Last updated: Mon Oct 5 15:20:46 2026
Sail support is experimental. It is only available in the development version of pysparklyr.
pak::pak("mlverse/pysparklyr")Intro
Sail is a Spark Connect server written in Rust. It does not need Java or a JVM. Sail runs the same Spark DataFrame and SQL commands that Spark does, so sparklyr can talk to it the same way it talks to Spark Connect. You can connect to a remote Sail server that is already running, or to a local Sail server on your machine.
Get started
To start and work with a local Sail server, use "sail" as the method. If you can use uv on your machine, sparklyr installs the Python libraries that Sail needs when you connect. It then starts a local Sail server and connects you to it. You do not need to do any other setup (a.k.a. no Java needed!).
The server uses a free port on your machine. spark_disconnect() stops the server. The server also stops when the R session ends.
Python environment
If you cannot use uv, run install_sail() once, before you connect. It creates a Python environment that stays on your machine. These are the main libraries in the new Python environment:
-
pyspark-client, which is a thin version ofpyspark -
pysail, whichsparklyruses to start local Sail servers -
rpy2, which runs the R code thatspark_apply()sends
pysparklyr::install_sail()After that, connect with spark_connect() as shown in Get started. sparklyr finds and uses the new environment, instead of having uv create a temporary one.
To install a specific Sail version, pass it in version:
pysparklyr::install_sail(version = "0.7")Work with data
Sail works with the same sparklyr and dplyr code that you use with Spark. Copy a local data frame to Sail with copy_to():
Then use dplyr verbs. sparklyr turns them into Spark SQL, and Sail runs the query:
Use collect() to bring the results into R:
mtcars_tbl |>
filter(hp > 150) |>
select(mpg, cyl, hp) |>
collect()
#> # A tibble: 13 × 3
#> mpg cyl hp
#> <dbl> <dbl> <dbl>
#> 1 18.7 8 175
#> 2 14.3 8 245
#> 3 16.4 8 180
#> 4 17.3 8 180
#> 5 15.2 8 180
#> 6 10.4 8 205
#> 7 10.4 8 215
#> 8 14.7 8 230
#> 9 13.3 8 245
#> 10 19.2 8 175
#> 11 15.8 8 264
#> 12 19.7 6 175
#> 13 15 8 335You can also run SQL directly with DBI:
DBI::dbGetQuery(sc, "SELECT cyl, COUNT(*) AS n FROM mtcars GROUP BY cyl")
#> cyl n
#> 1 6 7
#> 2 4 11
#> 3 8 14Run R code
spark_apply() runs an R function on each group of the data. Here, nrow() counts the rows for each value of am. columns sets the names and types of the result:
mtcars_tbl |>
spark_apply(nrow, group_by = "am", columns = "am double, x long")
#> # A query: ?? x 2
#> # Database: connect_sail
#> am x
#> <dbl> <dbl>
#> 1 1 13
#> 2 0 19With a "local" connection, the R code runs inside your R session, not in a separate R process. This means:
- If the R code crashes, it can end your R session
- The R code can change objects in your global environment
Connect to a running Sail server
You can also connect to a Sail server that is already running. Pass the server’s address in master. The address uses the “sc://” protocol. Your machine still needs the Python environment from Get started or Python environment:
sc <- spark_connect(
master = "sc://localhost:50051",
method = "sail"
)The Python environment that runs the Sail server needs pyspark-client. To use spark_apply(), the server also needs:
rpy2- R
- The same Python version that your R session uses
See the Sail documentation to learn how to start a Sail server.
How it works
sparklyr uses reticulate to call the Python pyspark-client library. That library sends the commands to the Sail server over gRPC. With a "local" connection, pysail runs the Sail server, which is written in Rust, inside your R session (Figure 1). With a remote connection, the commands go to a Sail server on another machine instead.
flowchart LR
subgraph rs[R session]
subgraph r[R]
sr[sparklyr]
rt[reticulate]
end
subgraph ps[Python]
pc[pyspark-client]
subgraph rust[Rust]
sl[Sail server]
end
end
end
sr <--> rt
rt <--> pc
pc <-- gRPC --> sl
style rs fill:#fff,stroke:#666,color:#000
style r fill:#fff,stroke:#666,color:#000
style sr fill:#fff,stroke:#666,color:#000
style rt fill:#fff,stroke:#666,color:#000
style ps fill:#fff,stroke:#666,color:#000
style pc fill:#fff,stroke:#666,color:#000
style rust fill:#fff,stroke:#666,color:#000
style sl fill:#fff,stroke:#666,color:#000
sparklyr communicates with a local Sail server
Limitations
Sail does not support some sparklyr features:
-
Caching -
copy_to()and thespark_read_*()functions needmemory = FALSE. Withmemory = TRUE, they return an error.compute()also returns an error. -
Machine learning - The
ml_*()andft_*()functions do not work. -
Tuning -
tune_grid_spark()does not work.