Stratified Sampling in Tableau Prep | New in tableau 2023.3
Tableau Prep 2023.3 adds stratified sampling — here's why it makes prepping large datasets faster and far more representative.
- Stratified sampling takes your population, groups it by a chosen dimension, then randomly selects within each group so the sample is more representative
- In Tableau Prep, sampling only affects what you see while prepping — when you hit run, all records flow through to the output regardless of sample settings
- Without stratification, a rare category like loan defaults (14% of the data) can be under-represented in your sample, leaving you with too few rows to analyse
- To enable it, go to the input step, open Data Sample, choose Stratified, then pick the column to balance on (e.g. client ID stratified by default)
- Sampling keeps row counts down to avoid Tableau Prep performance and stability issues with multi-million row datasets while preserving a useful spread
Tableau Prep 2023.3 adds stratified sampling, which forces your working sample to reflect the true spread of a chosen column instead of a plain random cut that can under-represent rare categories. Tim explains why that matters for prepping large datasets accurately and without performance pain.
Tim uses a credit dataset where only 14% of records represent loan defaults. He shows what happens to that rare category under normal random sampling versus stratified sampling.
- What stratified sampling means 0:22
You take your full population, group it by a chosen dimension, then randomly sample within each group rather than across the whole dataset at once. This keeps the sample balanced across that dimension instead of just reflecting whatever proportions happen to exist overall.
- How sampling works in Tableau Prep 1:39
Sampling only controls what you see while building your flow — it's set on the input step under Data Sample, with two separate settings: how many rows to bring in, and the method used to select them. Crucially, when you hit run, Tableau Prep pushes every record from the full dataset through to the output regardless of your sample settings.
- The under-representation problem 2:36
With plain random sampling, a rare category can drop even further out of view — in Tim's example, a 14% default rate in the full data fell to 13% in a 100-row sample, leaving too few rows to meaningfully inspect that group. On a multi-million row dataset this gets worse, effectively blinding you to minority categories while you're prepping.
- Turning on stratified sampling 4:39
On the input step, open Data Sample and switch the method to Stratified, then choose the column to stratify by (Tim picks the default field rather than the default client ID option). Prep then balances the sample across the groups in that column — in the demo, a 50/50 split of defaulters versus non-defaulters instead of the skewed original ratio.
- Why sample instead of running everything 6:03
Sampling exists to keep row counts manageable while you work, since pushing very large volumes through Prep can trigger performance and stability issues. Stratified sampling lets you keep that row count down for speed while still preserving a spread that's actually useful to look at.
- Choose the column that matches your analysis 6:41
As a general tip, stratify on whichever column best represents the thing you're actually trying to analyse — for demographic analysis of defaults, for instance, you'd stratify on the demographic field rather than the default field itself, so that field gets a representative spread instead.
- Sampling settings never change what's written to the output — running the flow always processes the full dataset, sampling is purely a prep-time view.
- With standard random sampling, rare categories can end up even more under-represented in your sample than in the full population, which is easy to miss until you check the exact counts.
- Stratification can be applied across different columns, so pick the one that matches whatever you're actually trying to analyse rather than leaving it on the default field.
Reach for stratified sampling when you're prepping a large dataset and need to eyeball a minority category (fraud, churn, defaults, rare events) while building your flow, without waiting on or destabilising Prep by processing every row.
How this Rollup was made provenance & method
A Rollup is drafted by AI from the video's transcript, then reviewed and edited by Tim. Everything used to produce this one is listed below — the model, the exact prompt, and the source video — so the process is transparent and reproducible.
- Transcription
- On-device — NVIDIA Parakeet v3 for recent videos, OpenAI Whisper large-v3 for earlier ones. The transcript never leaves the machine or gets published.
- Drafting
- Claude Sonnet 5 in the cloud, from that transcript.
- Prompt
- The exact Rollup prompt (v2) — the full system prompt, unedited.
- Source video
- Watch on YouTube
- Drafted
- 5 July 2026 at 09:38
- Reviewed & edited
- 5 July 2026 at 09:41 · by Tim Ngwena
Model + prompt + video is everything you'd need to recreate a Rollup like this yourself. The one thing we don't share is the transcript.
Rights. The video and its transcript are the property of TN Media Ltd. Unauthorised use or download is prohibited. © TN Media Ltd.