dask - Scale Python workflows with Dask
Scale pandas, NumPy, and custom Python workflows beyond memory or across clusters using DataFrames, Arrays, Bags, Futures, and schedulers.
Tags
Updated: 2026-10-02Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Process larger-than-memory datasets
- Parallelize pandas operations
- Parallelize NumPy arrays
- Process unstructured data
- Submit dependency-aware tasks
- Configure execution schedulers
- Scale workloads across clusters
- Inspect distributed diagnostics
Inputs
- Datasets and file paths
- Custom Python functions
- Workflow task definitions
- Chunking configuration
- Cluster configuration
- Cloud provider credentials
Outputs
- Computed data results
- Parallel task results
- Distributed diagnostics
- Scheduler dashboard
Requirements
- Python 3.10 or later
- dask 2026.8.0
- pandas 2 or later for DataFrames
- PyArrow 16 or later for DataFrames
- s3fs for s3:// paths
- gcsfs for gs:// paths
- dask.distributed for cluster deployment
- Python 3.12 or later for current Zarr 3.4
