Steps
The six kinds of step, what each one takes and produces, and how they connect.
A pipeline is a graph of steps. Each step does one thing — fetch data, transform it, run a model, store the result — and passes its output to the next as a named dataframe.
Step logic is reusable while its configuration is not: the same training script can be used by a dozen pipelines, each supplying its own features, target and coefficients. That is what makes a pipeline worth building rather than a script worth copying.
The step types
| Step | Takes | Produces | Used for |
|---|---|---|---|
| Query | nothing | one dataframe | Fetching time-series data. Every pipeline starts with at least one. |
| Script | any number of dataframes | any number of dataframes | Transformations, feature engineering, calls to a third-party service — any Python you write. |
| Model | a dataframe | a dataframe | Training a machine learning model, running inference with it, or both. |
| Write | a dataframe | nothing | Storing results back into the platform as datapoints. |
| Plot | a dataframe | nothing | Producing a Plotly figure to look at. |
| Alert | a dataframe | nothing | Raising a notification — push, email, SMS, WhatsApp — when a condition holds. |
Data passing between steps is not stored. If you want to keep a result, a Write step puts it in the platform; if you want to look at one, a Plot step draws it. Anything else is gone when the run ends.
How steps connect
Steps are wired by naming dataframes: a step declares the names it produces, and a later step declares the names it consumes. The platform checks the result is a valid graph before it will save the pipeline:
- Something must start it. At least one step with no inputs, which in practice is a Query.
- Every input must be satisfied by an output produced earlier in the graph.
- Output names must be unique. No two steps may produce a dataframe of the same name.
- One output can feed several steps — a single query can be plotted, written and alerted on — but each input is connected to exactly one output.
- No orphans and no cycles. Every step must be reachable, and the graph cannot loop back on itself.
A Query starts a branch and a Write, Plot or Alert ends one. Script, Model and Plot steps are flexible about how many dataframes they take and return.
Order is not position. Steps do not run in the order you added them. The engine works out the order from the graph: a step runs as soon as the dataframes it needs exist, and independent branches run in parallel.
What changes between an experiment and production
Two step types behave differently depending on how the pipeline is run, which is why a pipeline is normally tried as an experiment first and only then deployed:
| Step | In an experiment | In production |
|---|---|---|
| Model | Trains the model if a training script is configured, then runs inference. | Inference only — it uses the model already trained. |
| Plot | Generates the figure. | Skipped. |
The other four behave identically in both. See Pipeline for the difference between the two execution modes, and Experiment for running one.
Building them
Steps can be composed visually in the Pipeline Editor, by dragging blocks from the Library panel and connecting their outputs to inputs, or defined programmatically in Python or JSON. Working with pipelines covers both, and the anomaly detection tutorial builds a pipeline with five of these step types from end to end.
Updated 22 days ago