{pipeflow} is a lightweight framework for fast, interactive data analysis pipelines. Two CRAN releases landed since the introduction post: v0.3.0 rebuilt the core around a C++-powered DAG and a new functional API, and v0.4.0 added a more expressive interface. This post provides an overview of all recent changes and highlights one of the new features: pipeline views.
{pipeflow} builds a pipeline by adding one R function per step with pip_add (see also the Get started article). In this post, we’ll work with the following toy pipeline:
load produces a sequenceclean doubles itfit sums it, andreport formats the resultpip <- pip_new("my-pip") |>
pip_add(
step = "load",
fun = \(n = 5) seq_len(n),
tags = c("io", "daily")
) |>
pip_add(
step = "clean",
fun = \(x = ~load) x * 2,
tags = c("io", "core")
) |>
pip_add(
step = "fit",
fun = \(x = ~clean) sum(x),
tags = c("model", "core")
) |>
pip_add(
step = "report",
fun = \(x = ~fit) paste("result:", x),
tags = "report"
)
Each step must have a unique name, and function parameters can refer to other
steps — for example, x = ~load takes the output of the load step.
We also assign tags to each step, which we will use to filter steps later.
Before we run it, let’s have a look:
pip
<pipeflow> my-pip (4 steps)
---------------------------
step params depends state tags
1: load n new io,daily
2: clean x load new io,core
3: fit x clean new model,core
4: report x fit new report
---------------------------
<ready> last run: never
We see one row per step, showing its parameters (params), the steps it
depends on (depends), its current state (all new before the first run),
and our tags. Once we run the pipeline, …
pip_run(pip)
info [2026-10-03 11:52:38.120 UTC]: Starting run of pipeflow 'my-pip'
info [2026-10-03 11:52:38.121 UTC]: Step 1/4 load
info [2026-10-03 11:52:38.122 UTC]: Step 2/4 clean
info [2026-10-03 11:52:38.129 UTC]: Step 3/4 fit
info [2026-10-03 11:52:38.134 UTC]: Step 4/4 report
info [2026-10-03 11:52:38.136 UTC]: Finished run of pipeflow 'my-pip'
pip
<pipeflow> my-pip (4 steps)
---------------------------
step params depends state out tags
1: load n done 1,2,3,4,5 io,daily
2: clean x load done 2, 4, 6, 8,10 io,core
3: fit x clean done 30 model,core
4: report x fit done result: 30 report
---------------------------
<ready> last run: 2026-10-03 13:52:38
… there is a new out column with the results, which can also be
accessed directly:
pip[["report", "out"]]
[1] "result: 30"
Real pipelines get long, and often you only care about a subset of steps: a topic, a stage, or the steps that produce the desired outputs. Views (new in v0.4.0) let you work on such a subset without copying anything. They reference the underlying pipeline, so every operation applied to it — running it, updating parameters, collecting output, … — writes through to the original pipeline, restricted to the steps the view covers.
pip_view() returns a view containing only the steps that match the given
filters:
pip_view(pip, tags = "core")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip_view(pip, tags = "core", state = "done")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
By default steps must match all filters (logical AND), while the values
within one filter are alternatives (OR). join = "union" keeps steps that
match any filter, and fixed = FALSE treats filter values as regular
expressions:
pip_view(pip, tags = "report", step = "clean", join = "union")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: report x fit done result: 30 report
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
pip_view(pip, step = "^f", fixed = FALSE)
<pipeflow_view> my-pip view (1 of 4 steps)
------------------------------------------
step params depends state out tags
1: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
[The extract operator (also new in v0.4.0) provides a data.table-like way of selecting steps and returns a view by default:
pip[c("load", "fit")]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: load n done 1,2,3,4,5 io,daily
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
Boolean filters are evaluated in the context of the step table.
pip[tags %like% "core"]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
step params depends state out tags
1: clean x load done 2, 4, 6, 8,10 io,core
2: fit x clean done 30 model,core
------------------------------------------
<ready> last run: 2026-10-03 13:52:38
Running a view executes the steps it covers together with any upstream
dependencies that are not up to date. The run log marks the view’s own steps
as [view] and the dependencies pulled in as [upstream]:
pip_reset(pip) # reset the pipeline to "new" state for demonstration
pip_run(pip_view(pip, step = "report"))
info [2026-10-03 11:52:38.298 UTC]: Starting run of pipeflow 'my-pip view'
info [2026-10-03 11:52:38.298 UTC]: Step 1/4 [upstream] load
info [2026-10-03 11:52:38.298 UTC]: Step 2/4 [upstream] clean
info [2026-10-03 11:52:38.300 UTC]: Step 3/4 [upstream] fit
info [2026-10-03 11:52:38.301 UTC]: Step 4/4 [view] report
info [2026-10-03 11:52:38.302 UTC]: Finished run of pipeflow 'my-pip view'
Afterwards, the original pipeline is up to date for the covered steps. Let’s collect all results related to the “core” steps:
pip[tags %like% "core"] |> pip_collect()
$clean
[1] 2 4 6 8 10
$fit
[1] 30
For a complete guide on composing views, see the Pipeline views vignette.
Views are only one of the additions. Each of the following has its own article on the documentation site:
pip_*() functions,
with every function also available as a method (p$run(), p$view(), …):
Get started.[ and [[ for
reading and editing steps:
reference.pip_collect() with grouping by tags or
dependencies: Collect and group output.rbind(), or replace, rename, remove, and reset steps:
Combine pipelines ·
Modify existing pipelines.exec = "split", "auto", and
"reduce" modes: Split, map, and reduce..self, run-time restructuring, and
restart(): Self-modifying pipelines.Under the hood, significant performance gains have been made by implementing the
dependency graph in C++ (since v0.3.0) and by improving the step and parameter
bookkeeping in the underlying {data.table}.
Finally, by dropping lgr and jsonlite, external package dependencies have
been reduced to {data.table} and Rcpp.
{targets} remains R’s de-facto standard for heavy-duty, reproducible
pipelines: it persists results to disk, records provenance, and scales to
distributed infrastructure via crew. {pipeflow} does not try to compete on
that home turf but targets a different niche — the interactive session,
often as the backend of a Shiny app — where a pipeline is built and modified
on the fly and has to respond while a user waits.
Since responsiveness is key in that use case, {pipeflow} has been optimized
for low latency.
The pipeflow vs targets article compares the in-session latency of the two packages head to head. The figure below gives a flavour of the results:

It shows the scenario where one parameter has changed1 and the pipeline is rerun, with each step doing 5 ms of work. The x-axis represents the number of steps in the pipeline (its size), and the y-axis shows the time required to re-run it.2
Across the tested sizes, {pipeflow} is consistently the faster of the two — see the article for the full benchmark and methodology.
Overall, the two are best seen as complementary: reach for {targets} for large batch jobs that must be reproducible end to end, and for {pipeflow} when the pipeline is part of an interactive application.
{pipeflow} 0.4.0 is on CRAN:
install.packages("pipeflow")
The full changelog lists every change, and the documentation includes the articles linked above. If you build Shiny apps or spend a lot of time exploring parameters interactively, give it a try — feedback and issues are welcome on GitHub.
Text and figures are licensed under Creative Commons Attribution CC BY-SA 4.0. The figures that have been reused from other sources don't fall under this license and can be recognized by a note in their caption: "Figure from ...".