dask - Scale Python data and compute workflows with Dask
Scales pandas, NumPy, and custom Python workflows across memory limits, cores, or clusters using Dask collections and schedulers.
Tags
Updated: 2026-10-03Capabilities
What this skill does
- Scale tabular data processing
- Process blocked array computations
- Process unstructured records
- Submit dependent tasks
- Select execution schedulers
- Monitor distributed execution
Inputs
- Tabular data files
- Scientific array datasets
- Text, JSON, or log files
- Python task functions
- Scheduler configuration
- Cluster resources
- Cloud storage credentials
Outputs
- Computed DataFrames
- Computed arrays and values
- Distributed task results
- Execution dashboard and diagnostics
Requirements
- Python 3.10 or later
- dask 2026.8.0
- pandas 2 or later for DataFrames
- PyArrow 16 or later for DataFrames
- s3fs or gcsfs for cloud paths
- dask.distributed for cluster deployment
