clip()
Clip all numeric columns to a physical range. Properties 'min' and/or 'max' (floats) bound the values; at least one of the two must be provided.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
clustering()
K-means clustering with automatic cluster count selection via silhouette score.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
diff()
Compute discrete differences (rate of change) across all numeric columns. Property 'periods' (default 1) is the number of rows to shift before subtracting — use 1 for first derivative, higher for longer horizons.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
fillna()
Fill missing values per column using the 'fillna' property dict mapping column names to fill values.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
filter_df()
Filter dataframe rows using pandas query strings from the 'filters' property list.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
formula()
Add a new column computed from a user-defined math expression over existing columns. Property 'expression' is required — intuitive math syntax referencing column names (e.g. 'temp_1 + temp_2', '(p_in - p_out) / p_in * 100', 'sqrt(vibration_x2 + vibration_y2)', 'clip(temperature, 0, 500)', 'log(power + 1)'). Supports arithmetic operators (+, -, , /, %, *) and these functions: abs, sqrt, log, log10, exp, clip, round, min, max. Column names with spaces must be wrapped in backticks (e.g. 'motor temp * 2'). Property 'result' (default 'result') names the output column.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
interpolate()
Interpolate missing values in numeric columns. The 'method' property selects the pandas interpolation method: 'time' (default, respects timestamp spacing), 'linear', 'nearest', 'pad', 'polynomial', or 'spline'.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
lag()
Add lagged versions of all numeric columns as new '_lag
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
merge()
Merge multiple dataframes by index using outer join (configurable via 'how' property).
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
normalize()
Normalize all numeric columns using 'minmax' (default) or 'zscore' scaling (configurable via 'method' property).
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
pca()
Reduce numeric columns to principal components, replacing them with 'PC1', 'PC2', ... Property 'n_components' is required — an integer (number of components) or a float in (0, 1] (minimum explained variance ratio). NaN values must be handled upstream (e.g. with interpolate or fillna).
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
remove_outliers()
Drop rows containing outliers in any numeric column. The 'method' property selects 'iqr' (default, Tukey 1.5*IQR rule) or 'zscore' (drops rows beyond ±3 sigma).
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
resample()
Resample the dataframe to a new time frequency. Property 'freq' is required (pandas offset alias, e.g. '1min', '5min', '1H', '1D'); 'agg' selects the aggregation ('mean' default, 'sum', 'min', 'max', 'first', 'last', 'median').
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
rolling()
Apply a rolling-window aggregation over all numeric columns. Property 'window' is required (integer number of rows); 'agg' selects the aggregation ('mean' default, 'sum', 'std', 'min', 'max', 'median').
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
setpoint_deviation()
Add '
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
steady_state_filter()
Keep only rows where all numeric columns are in steady state (rolling std over 'window' rows stays below 'tolerance'). Drops transients and start-up periods — a standard preprocessing step before process modeling.
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |
target_feat_to_binary()
Convert a target feature column to binary (0/1) using a threshold operator (lt, lte, gt, gte).
| Name | Type | Default | Description |
|---|---|---|---|
| context | Context |