Skip to main content

Kafka

Apache Kafka

Pronunciation
KAHF-kuh
Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/kafka

In short

Apache Kafka is a distributed event streaming platform that stores events in durable, ordered logs for many services to publish, read in real time or replay.

What is Kafka?

Kafka was built at LinkedIn to move huge volumes of activity data and was open-sourced in 2011; it is now an Apache Software Foundation project. Producers write events, such as "order placed" or "page viewed", to named topics. Kafka appends each event to the end of a log and keeps it for a configured time, days or even forever, whether or not anyone has read it yet.

Each topic is split into partitions spread across a cluster of servers called brokers. Events within one partition keep their order, and each has a numbered position called an offset. Consumers read at their own pace and remember their offset, so a slow or restarted consumer simply continues where it left off, and a new one can replay history from the beginning.

Consumers that share a group name split the partitions between them, so adding consumers spreads the work, while separate groups each receive every event. This lets one stream of orders feed billing, shipping and analytics independently. Kafka is used for event-driven microservices, activity tracking, log and metric pipelines, and change data capture from databases.

A common misconception is that Kafka is just a message queue. A queue usually deletes a message once it is handled, while Kafka keeps the log and lets any number of readers go back in time. That power comes with operational weight: partitions, replication and retention need planning, so smaller systems often start with a simpler broker such as RabbitMQ.

At a glance

A producer writes order events to the topic "orders", which is split into three partitions P0, P1 and P2; each partition is an ordered log where new events are appended at the end and every event has an offset. In the consumer group "billing", consumer 1 reads P0 and P1 and consumer 2 reads P2. The separate group "analytics" has one consumer that reads every partition.Producerorder eventstopic: ordersP001234P10123P201234offsets 0, 1, 2… newest at the rightgroup: billingConsumer 1reads P0 and P1Consumer 2reads P2group: analyticsConsumerreads every partition
Inside one group, the partitions are shared out so the work is split; every group still gets every event, and each consumer keeps its own offset.

Key takeaways

  • Kafka stores events in durable, ordered, append-only logs called topics.
  • Topics are split into partitions spread across a cluster of brokers.
  • Consumers track their own offset, so they can resume or replay history.
  • Consumer groups share partitions; separate groups each get every event.
  • It is more powerful, and heavier to run, than a simple message queue.

Example

Producing and consuming events (Node.js with kafkajs)javascript
import { Kafka } from "kafkajs";

const kafka = new Kafka({ brokers: ["localhost:9092"] });

// Producer: append an event to the "orders" topic
const producer = kafka.producer();
await producer.connect();
await producer.send({ topic: "orders", messages: [{ key: "1001", value: '{"total": 49}' }] });

// Consumer: members of "billing" share the topic's partitions
const consumer = kafka.consumer({ groupId: "billing" });
await consumer.connect();
await consumer.subscribe({ topic: "orders", fromBeginning: true });
await consumer.run({ eachMessage: async ({ message }) => console.log(message.value.toString()) });

Readers ask

Is Kafka a message queue?

Not exactly. It can be used like one, but Kafka keeps events in a log after they are read, so many consumers can read the same events independently and replay them later. A classic queue removes a message once it is handled.

What is a Kafka topic?

A topic is a named stream of events, such as "orders". It is split into partitions, which are ordered logs stored across the brokers of the cluster.

Does Kafka still need ZooKeeper?

No. Newer versions manage the cluster themselves with a built-in mode called KRaft, and Kafka 4.0 removed ZooKeeper support entirely.

Often compared

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings