An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

2.5.2. Format Selection and Conversion Patterns

💡 First Principle: Format conversion should happen as early in the pipeline as possible — the earlier you convert to an efficient format, the more downstream processes benefit. The raw landing zone stores data in whatever format it arrives (CSV, JSON), but the curated zone should always use a columnar, compressed format.

The common data lake pattern uses three zones:

Raw zone: Data lands in its original format. No transformation. This preserves the original data for reprocessing and auditing.

Curated zone: Data is cleaned, typed, and converted to Parquet or ORC with compression. Partitioned by commonly filtered columns (date, region, product category). This is where Glue ETL or EMR Spark jobs do the heavy lifting.

Serving zone: Data is further aggregated or denormalized for specific consumption patterns — dashboards, ML feature stores, or API responses. May live in Redshift, DynamoDB, or a dedicated S3 prefix.

Zone access and naming. Access widens as data is refined: the raw zone is restricted to ingestion/ETL roles, curated to data engineers and approved analysts, and serving is broadly readable by consumers (enforced per prefix with IAM/bucket policies or Lake Formation grants). The same three layers are often called bronze / silver / gold (the medallion architecture): bronze = raw, silver = curated, gold = serving.

Conversion services: Glue ETL jobs are the primary tool for format conversion. A typical Glue job reads CSV/JSON from the raw zone, applies transformations (data type casting, null handling, deduplication), writes Parquet to the curated zone with Snappy compression, and partitions by date. Firehose can also convert to Parquet on delivery using its built-in format conversion feature — simpler for streaming ingestion. For a one-off or SQL-driven conversion, Athena CREATE TABLE AS SELECT with format = 'PARQUET' writes the converted copy serverlessly (syntax in 2.7.1).

Partition and file-size rules. Partition on low-cardinality columns that queries actually filter on (date, region, a 10-value device type). A high-cardinality key such as customer_id or device_id explodes into huge numbers of tiny partitions and files, and the catalog and per-file overhead outweigh any pruning. File size matters too: each file costs an S3 GET, a footer/metadata read and reader start-up, so 100,000 files of 100 KB query far slower than the same data in files of roughly 128 MB or more. Fix small files by compacting them (Glue job, Athena CTAS, or Iceberg compaction), not with more query concurrency (each query still opens every tiny file) or finer partitions (which only split the data into even smaller files), and prevent them upstream with larger Firehose buffers.

⚠️ Exam Trap: Firehose's built-in Parquet conversion requires a Glue Data Catalog table to define the target schema. If the question mentions Firehose delivering Parquet to S3, a Glue table must exist. This is a common trick question — candidates forget the Glue dependency.

Reflection Question: Why should the raw zone keep data in its original format even though columnar formats are more efficient? In what scenario would you need to reprocess from the raw zone?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications