An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

2.1.2. Amazon Kinesis Data Firehose

💡 First Principle: Firehose is the "load" service for streaming data — it captures, optionally transforms, and delivers data to a destination. If Kinesis Data Streams is the highway, Firehose is the delivery truck that picks up from the highway and drops off at S3, Redshift, OpenSearch, or an HTTP endpoint.

Firehose eliminates the need to write consumer code. You configure a source (direct PUT, Kinesis Data Streams, or MSK), an optional transformation (Lambda function), and a destination — and Firehose handles buffering, batching, compression, encryption, and retry logic automatically.

The critical distinction: Firehose is near-real-time, not real-time. It buffers incoming records and delivers when either the buffer size (1–128 MB) or the buffer interval (0–900 seconds, default 300) is reached — whichever comes first. A zero-second interval cuts delivery to a few seconds, but Firehose still only delivers batches; it doesn't process individual events. When an exam question requires sub-second processing, Firehose alone is insufficient — you need KDS or MSK with a custom consumer or Managed Flink.

Transformation with Lambda. Firehose can invoke a Lambda function to transform each batch of records before delivery. Common uses: enriching records with lookup data, filtering irrelevant events, reformatting timestamps, or turning CSV and other non-JSON input into JSON. The Lambda function receives a batch of records, processes them, and returns results with a status per record (Ok, Dropped, or ProcessingFailed). Keep the function lean: cache slowly changing lookup data (such as a DynamoDB reference table) in the execution environment, outside the handler, instead of calling the database for every record. Plain JSON→Parquet/ORC needs no Lambda at all — Firehose's built-in record format conversion does it, using a Glue Data Catalog table as the target schema (see 2.5.2).

Delivery destinations: S3 (most common), Amazon Redshift (via S3 staging + COPY), Amazon OpenSearch Service, HTTP endpoints (Splunk, Datadog, custom APIs), and third-party destinations. For Redshift delivery, Firehose first writes to S3, then issues a COPY command — understanding this two-step process is exam-relevant.

Error handling. Failed records can be sent to a separate S3 prefix (backup bucket), ensuring no data loss even when the Lambda transform or destination fails. For OpenSearch, Splunk and HTTP endpoint destinations, the S3 backup can be set to all records, so one stream feeds the search index and archives everything in S3.

⚠️ Exam Trap: When a question says "near-real-time delivery to S3" with "minimal operational overhead," Firehose is almost always the answer — not a Lambda consumer writing to S3 (more code to maintain) or a Glue streaming job (more complex). But if the question says "real-time processing with sub-second latency," you need Kinesis Data Streams with a custom consumer or Managed Apache Flink.

Reflection Question: A company wants to stream application logs to S3 in Parquet format for Athena queries. They want the simplest solution with no custom consumer code. What's the architecture?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications