Skip to main content

Create a new pipeline

Each pipeline is represented by a YAML file in a folder named pipelines/ under the Mage project directory. For example, if your project is named demo_project and your pipeline is named etl_demo then you’ll have a folder structure that looks like this:
Create a new folder in the demo_project/pipelines/ directory. Name this new folder after the name of your pipeline. Add 2 files in this new folder:
  1. __init__.py
  2. metadata.yaml
In the metadata.yaml file, add the following content:
Change etl_demo to whatever name you’re using for your new pipeline.

Sample pipeline metadata content

This sample pipeline metadata.yaml will produce the following block dependencies:
Sample pipeline

metadata.yaml sections

Pipeline attributes

array of objects
An array of blocks that are in the pipeline.
string
Unique name of the pipeline.
string enum
The type of pipeline. Currently available options are:
  • databricks
  • integration
  • pyspark
  • python (most common)
  • streaming
string
Unique identifier of the pipeline. This UUID must be unique across all pipelines.
string
Optional description of what the pipeline does.
string enum
Pipeline level executor type. Supported values:
  • ecs
  • gcp_cloud_run
  • azure_container_instance
  • k8s
  • local_python (most common)
  • pyspark
integer
Number of concurrent executors to run the pipeline. Used in streaming pipeline.
object
Optional configuration specific to the selected executor type. Refer to the following documentation for executor-specific options:
object
Spark-specific configuration for PySpark pipelines. Mirrors the keys you would normally pass to SparkConf (e.g., spark_master, executor_env, spark_jars).
object
Retry configuration at the pipeline level. See documentation for details.
object
Configuration for pipeline notification messages (e.g., on failure or success). See documentation for details.
object
Concurrency settings for block execution within the pipeline. See documentation for details.
  • block_run_limit: Maximum number of blocks that can run in parallel.
  • pipeline_run_limit
  • pipeline_run_limit_all_triggers
  • on_pipeline_run_limit_reached
boolean
Whether to cache block output in memory during execution.
boolean
If true, runs all blocks in a single process or k8s pod.
array of objects
Blocks that run after your main graph finishes (e.g., notifications, cleanup). Same shape as items in blocks.
array of objects
Conditional blocks that can short-circuit or branch execution. Same shape as items in blocks.
array of objects
Extension blocks (e.g., Great Expectations). Same shape as items in blocks.
array of objects
Chart or reporting blocks shown in the pipeline UI. Same shape as items in blocks.
object
Pipeline-level settings; currently supports triggers.save_in_code_automatically to control whether new/updated triggers are persisted to YAML.
object
Key/value pairs available to the pipeline for templating. Useful for environment-specific values that are not secrets.
object
Configure an external state store for pipeline runs.
array of strings
Free-form labels to organize and search for pipelines.
string
Identifier for the user or process that created the pipeline.
object
Environment-specific overrides for any top-level pipeline field. Mage Pro only. See environment overrides for guidance.

Block attributes

array of strings
An array of block UUIDs that depend on this current block. These downstream blocks will have access to this current block’s data output.
string enum
The method for running this block of code. Currently available options are:
  • ecs
  • gcp_cloud_run
  • azure_container_instance
  • k8s
  • local_python (most common)
  • pyspark
object
Optional configuration specific to the selected executor type. Refer to the following documentation for executor-specific options:
string enum
Programming language used by the block. Supported values:
  • python (most common)
  • r
  • sql
  • yaml
string
Unique name of the block.
string enum
The type of block. Currently available options are:
  • chart
  • custom (most common)
  • data_exporter
  • data_loader
  • dbt
  • scratchpad
  • sensor
  • transformer
The type of block will determine which folder it needs to be in. For example, if the block type is data_loader, then the file must be in the [project_name]/data_loaders/ folder. It can be nested in any number of subfolders.
array of strings
An array of block UUIDs that this current block depends on. These upstream blocks will pass its data output to this current block.
string
Unique identifier of the block. This UUID must be unique within the current pipeline. The UUID corresponds to the name of the file for this block.For example, if the UUID is load_data and the language is python, then the file name will be load_data.py.
string
Optional HEX color used to visually organize blocks in the UI.
object
Block-specific configuration. For integrations or dbt blocks this mirrors the configuration visible in the block settings panel.
string
Markdown content rendered in the block’s documentation tab.
integer
Optional priority for scheduling blocks when multiple are runnable.
object
Retry configuration at the block level. See documentation for details.
array of strings
Labels applied to the block for organization and filtering.
integer
Maximum execution time for the block in seconds.