pipeflow: v0.4.0 on CRAN

pipeflow R package pipeline tools

{pipeflow} is a lightweight framework for fast, interactive data analysis pipelines. Two CRAN releases landed since the introduction post: v0.3.0 rebuilt the core around a C++-powered DAG and a new functional API, and v0.4.0 added a more expressive interface. This post provides an overview of all recent changes and highlights one of the new features: pipeline views.

Pipeline setup

{pipeflow} builds a pipeline by adding one R function per step with pip_add (see also the Get started article). In this post, we’ll work with the following toy pipeline:

pip <- pip_new("my-pip") |>
    pip_add(
        step = "load",
        fun = \(n = 5) seq_len(n),
        tags = c("io", "daily")
    ) |>
    pip_add(
        step = "clean",
        fun = \(x = ~load) x * 2,
        tags = c("io", "core")
    ) |>
    pip_add(
        step = "fit",
        fun = \(x = ~clean) sum(x),
        tags = c("model", "core")
    ) |>
    pip_add(
        step = "report",
        fun = \(x = ~fit) paste("result:", x),
        tags = "report"
    )

Each step must have a unique name, and function parameters can refer to other steps — for example, x = ~load takes the output of the load step. We also assign tags to each step, which we will use to filter steps later.

Before we run it, let’s have a look:

pip
<pipeflow> my-pip (4 steps)
---------------------------
     step params depends state       tags
1:   load      n           new   io,daily
2:  clean      x    load   new    io,core
3:    fit      x   clean   new model,core
4: report      x     fit   new     report
---------------------------
<ready> last run: never

We see one row per step, showing its parameters (params), the steps it depends on (depends), its current state (all new before the first run), and our tags. Once we run the pipeline, …

pip_run(pip)
info [2026-10-03 11:52:38.120 UTC]: Starting run of pipeflow 'my-pip'
info [2026-10-03 11:52:38.121 UTC]: Step 1/4 load
info [2026-10-03 11:52:38.122 UTC]: Step 2/4 clean
info [2026-10-03 11:52:38.129 UTC]: Step 3/4 fit
info [2026-10-03 11:52:38.134 UTC]: Step 4/4 report
info [2026-10-03 11:52:38.136 UTC]: Finished run of pipeflow 'my-pip'
pip
<pipeflow> my-pip (4 steps)
---------------------------
     step params depends state            out       tags
1:   load      n          done      1,2,3,4,5   io,daily
2:  clean      x    load  done  2, 4, 6, 8,10    io,core
3:    fit      x   clean  done             30 model,core
4: report      x     fit  done     result: 30     report
---------------------------
<ready> last run: 2026-10-03 13:52:38

… there is a new out column with the results, which can also be accessed directly:

pip[["report", "out"]]
[1] "result: 30"

Pipeline views in action

Real pipelines get long, and often you only care about a subset of steps: a topic, a stage, or the steps that produce the desired outputs. Views (new in v0.4.0) let you work on such a subset without copying anything. They reference the underlying pipeline, so every operation applied to it — running it, updating parameters, collecting output, … — writes through to the original pipeline, restricted to the steps the view covers.

Creating views

pip_view() returns a view containing only the steps that match the given filters:

pip_view(pip, tags = "core")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip_view(pip, tags = "core", state = "done")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip_view(pip, step = c("clean", "fit"))
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38

By default steps must match all filters (logical AND), while the values within one filter are alternatives (OR). join = "union" keeps steps that match any filter, and fixed = FALSE treats filter values as regular expressions:

pip_view(pip, tags = "report", step = "clean", join = "union")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
     step params depends state            out    tags
1:  clean      x    load  done  2, 4, 6, 8,10 io,core
2: report      x     fit  done     result: 30  report
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip_view(pip, step = "^f", fixed = FALSE)
<pipeflow_view> my-pip view (1 of 4 steps)
------------------------------------------
   step params depends state out       tags
1:  fit      x   clean  done  30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38

Selecting steps with [

The extract operator (also new in v0.4.0) provides a data.table-like way of selecting steps and returns a view by default:

pip[c("load", "fit")]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
   step params depends state       out       tags
1: load      n          done 1,2,3,4,5   io,daily
2:  fit      x   clean  done        30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38

Boolean filters are evaluated in the context of the step table.

pip[tags %like% "core"]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip[step %in% c("clean", "fit") & state == "done"]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38

Running views

Running a view executes the steps it covers together with any upstream dependencies that are not up to date. The run log marks the view’s own steps as [view] and the dependencies pulled in as [upstream]:

pip_reset(pip)  # reset the pipeline to "new" state for demonstration

pip_run(pip_view(pip, step = "report"))
info [2026-10-03 11:52:38.298 UTC]: Starting run of pipeflow 'my-pip view'
info [2026-10-03 11:52:38.298 UTC]: Step 1/4 [upstream] load
info [2026-10-03 11:52:38.298 UTC]: Step 2/4 [upstream] clean
info [2026-10-03 11:52:38.300 UTC]: Step 3/4 [upstream] fit
info [2026-10-03 11:52:38.301 UTC]: Step 4/4 [view] report
info [2026-10-03 11:52:38.302 UTC]: Finished run of pipeflow 'my-pip view'

Afterwards, the original pipeline is up to date for the covered steps. Let’s collect all results related to the “core” steps:

pip[tags %like% "core"] |> pip_collect()
$clean
[1]  2  4  6  8 10

$fit
[1] 30

For a complete guide on composing views, see the Pipeline views vignette.

More new and updated features

Views are only one of the additions. Each of the following has its own article on the documentation site:

Under the hood, significant performance gains have been made by implementing the dependency graph in C++ (since v0.3.0) and by improving the step and parameter bookkeeping in the underlying {data.table}. Finally, by dropping lgr and jsonlite, external package dependencies have been reduced to {data.table} and Rcpp.

{pipeflow} vs {targets}

{targets} remains R’s de-facto standard for heavy-duty, reproducible pipelines: it persists results to disk, records provenance, and scales to distributed infrastructure via crew. {pipeflow} does not try to compete on that home turf but targets a different niche — the interactive session, often as the backend of a Shiny app — where a pipeline is built and modified on the fly and has to respond while a user waits. Since responsiveness is key in that use case, {pipeflow} has been optimized for low latency.

The pipeflow vs targets article compares the in-session latency of the two packages head to head. The figure below gives a flavour of the results:

Warm incremental update with 5 ms of work per step for pipeflow and targets

It shows the scenario where one parameter has changed1 and the pipeline is rerun, with each step doing 5 ms of work. The x-axis represents the number of steps in the pipeline (its size), and the y-axis shows the time required to re-run it.2

Across the tested sizes, {pipeflow} is consistently the faster of the two — see the article for the full benchmark and methodology.

Overall, the two are best seen as complementary: reach for {targets} for large batch jobs that must be reproducible end to end, and for {pipeflow} when the pipeline is part of an interactive application.

Wrapping up

{pipeflow} 0.4.0 is on CRAN:

install.packages("pipeflow")

The full changelog lists every change, and the documentation includes the articles linked above. If you build Shiny apps or spend a lot of time exploring parameters interactively, give it a try — feedback and issues are welcome on GitHub.

Reuse

Text and figures are licensed under Creative Commons Attribution CC BY-SA 4.0. The figures that have been reused from other sources don't fall under this license and can be recognized by a note in their caption: "Figure from ...".