Table of Contents
Table of Contents
A Redpanda broker is being replaced. At the same time, the consumer-lag dashboard turns red. The first instinct is often to treat the two events as one failure: the partition moved, so the consumer group must have rebalanced. That shortcut can send an on-call engineer toward the wrong control—adding consumers while the real delay is a leader transition, or restarting a stable group while a slow handler is still holding it back.
A Kafka consumer group and a broker's partition ownership have related timelines, but they are not the same timeline. A consumer group assigns partitions to members. A partition reassignment moves replicas or leadership between brokers. Lag is the observable gap left when the consumer cannot advance as fast as records become available. Keeping those three facts separate is the fastest way to diagnose a Redpanda incident without guessing.
1What consumer lag hides
For a partition, lag is a distance between the log end and a group's committed progress. The exact display depends on the tool: one view may show committed offsets, another may show a consumer position, and an application metric may report records processed after downstream work completes. These measurements answer different questions. A group can fetch records quickly while committing slowly, or commit before a separate sink has finished its own work.
That distinction matters when a broker changes. A consumer may temporarily receive fetch errors while the partition leader moves, then recover with the same assignment. A member may leave the group because a restart or an application pause prevents it from meeting its client deadlines, which can cause another assignment. A partition may remain assigned throughout, yet its lag grows because the handler or a downstream service slowed down. The lag line alone cannot tell these stories apart.
Use a short evidence window around the first rise in lag. Record the partition, current and log-end offsets, active member, group state, fetch errors, and application processing time at the same time. In Redpanda, rpk group describe is useful for inspecting group membership and offsets; the fields and flags available to you depend on the Redpanda version and the way your cluster connection is configured. The consumer-offsets documentation explains the offset model, while the Apache Kafka consumer documentation provides the protocol context.
A useful first cut is this:
| Observation | What it proves | What it does not prove |
|---|---|---|
| Lag rises on one partition | Work is accumulating unevenly | That the consumer group rebalanced |
| Fetch errors line up with a leader move | The read path was interrupted or retried | That offsets were lost |
| Member count changes with lag growth | Group membership changed near the incident | That the broker move caused the membership change |
| Group is stable but handler time rises | Application completion slowed | That broker capacity is the bottleneck |
| Lag falls while no assignment change occurs | The group is draining backlog | That the original cause was storage |
The operational question is therefore not “How do we make lag smaller?” It is “Which timeline stopped making progress, and which timeline continued?”
2Assignment and partition movement are different operations
Kafka group coordination starts with members joining a group, subscribing, and receiving an assignment from the group protocol. The assignment strategy determines how subscribed partitions are distributed among members. A broker-side reassignment does something else: it changes where a partition's replicas are stored or which broker is serving leadership. The consumer still identifies the partition by topic and partition number; it does not receive a different logical partition because the replica moved.
That separation is often lost during a Redpanda maintenance window. Redpanda provides commands for moving partitions, balancing a cluster, decommissioning brokers, and performing rolling restarts. Those operations can change leaders and fetch paths without requiring a consumer-group rebalance. A group rebalance can happen at the same time if members restart, exceed their poll or heartbeat deadlines, or otherwise leave and rejoin the group, but it is a separate event to verify.
Think about the timelines as three lanes:
Broker lane: replica move → leader change → fetch path recovers
Group lane: members join/leave → assignment protocol → stable assignment
Work lane: records available → records fetched → records processed/committed
The lanes interact, but one lane does not imply a state transition in the others. A leader change can briefly slow the work lane while the group lane remains stable. A rolling deployment can trigger a group rebalance while partition leadership stays where it was. A large backlog can keep lag high after both lanes have returned to normal.
This is why an incident review should include both a broker event timeline and a group event timeline. For broker-side evidence, use the partition-move command reference, partition balancing guidance, and rolling-restart guidance. For group-side evidence, use the group description command and your consumer client's rebalance, assignment, and error logs.
3Three ways a broker change can surface as lag
3.1A leader transition interrupts fetches
When a partition's leader changes, clients must discover the target leader and retry fetches. The pause can be brief or prolonged, depending on the client, network path, broker health, and whether the target leader is ready to serve reads. During that window, the log end can continue advancing while the consumer's committed progress does not.
The signature is a narrow correlation: fetch errors or metadata refreshes appear close to the leader event, and the affected partitions recover without a membership change. Check the consumer's error log, broker request metrics, and the partition's leader history. If the group remains stable and handler time is unchanged, adding a consumer will not address the interruption. Wait for the leader path to recover, or follow the Redpanda recovery procedure if the partition remains unavailable.
3.2Membership churn turns maintenance into a rebalance
A broker replacement often comes with a deployment, pod eviction, or process restart. If the consumer process is colocated with the broker, shares a node pool, or is rolled by the same automation, the group may see a member leave. A member can also leave when application processing blocks polling long enough to violate the client configuration. Kafka then runs its group protocol again, and work may pause or be redistributed.
Look for a change in member IDs, assignment logs, and group state before changing capacity. The Kafka group configuration reference and the selected partition-assignment strategies explain which client settings are involved; exact defaults and supported protocols vary by client and version. In Redpanda, verify the group state with rpk group describe and compare it with the client logs rather than inferring a rebalance from lag alone.
If membership is unstable, stabilize the process lifecycle first. Check readiness and termination handling, container eviction, network reachability to the group coordinator, and the time spent between polls. Increasing the number of members before the group can stay stable usually adds another join and leave cycle to the incident.
3.3Data movement competes with the work path
Partition movement can consume broker resources and network bandwidth. The impact depends on the workload, replication settings, throttles, storage, and the amount of data that must be copied or recovered. A consumer may continue to fetch, but with higher request latency; or the broker may prioritize recovery work while the consumer falls behind.
The useful comparison is not “reassignment is running” versus “reassignment is not running.” Compare consumer fetch latency and throughput with the partition-move status, broker network and disk activity, and the lag slope. If fetch latency rises only for partitions involved in the move, the movement is a credible contributor. If handler time rises on every partition while broker request time is flat, investigate the application or downstream dependency instead.
Redpanda's continuous data-balancing documentation and broker decommissioning guidance describe controls for cluster movement. Treat the documented controls and their version-specific behavior as the source of truth. Do not assume a move is harmless because the group protocol did not run, and do not assume the group protocol ran because lag increased.
4A broker-change runbook for latency-sensitive consumers
The goal of a runbook is to preserve evidence while the system is changing. Capture the state before you remediate so that a recovered graph does not erase the cause.
4.1Before the change
Define which consumers must maintain freshness during the maintenance window and which partitions they read. Confirm that each group's committed-offset view, application processing rate, and partition-level lag are visible. Record the current group state and member set with the tooling used by your Redpanda deployment.
Separate broker work from consumer deployment work when possible. If both are needed, schedule them as distinct change steps so a member departure can be attributed to the right action. Make sure graceful shutdown allows the client to leave the group cleanly, and confirm that your orchestrator will not evict all members of a latency-sensitive group at once.
For a planned partition move, record the affected topics and partitions, the source and target brokers, and the command or controller action that initiated it. Redpanda's rpk cluster partitions move-status command can be used with the corresponding connection settings to follow a move; use the versioned command reference for the exact invocation.
4.2During the change
Watch the three lanes together. A compact incident panel should include:
- Group state and membership: stable, preparing, completing, empty, or another state exposed by the client and broker tooling. The names are protocol and implementation states, so use the values emitted by the deployed versions.
- Assignment changes: member assignment logs, revoked and assigned partitions, and the time from the first revoke to usable assignment.
- Partition movement: move status, leader changes, replica health, and broker request or network saturation for the affected partitions.
- Work progress: log-end offset, committed offset, lag slope, records consumed, handler duration, and downstream error or retry rate.
An alert should combine a lag condition with a state or rate change. A fixed lag threshold can be useful as a paging guardrail, but it is not a universal answer: the same offset gap may be harmless for a batch consumer and urgent for a low-latency pipeline. Alert on the slope of lag, time without committed progress, repeated group transitions, and the freshness objective that the application has declared.
Avoid resetting offsets during a broker change unless the incident plan explicitly calls for it. A lag spike caused by a leader transition or temporary rebalance is not evidence that records are corrupt. Offset resets change the application's position and can turn a recoverable delay into replay or data loss.
4.3After the change
Wait for both recovery conditions: the broker operation is complete, and the consumer has resumed useful work. A group can return to a stable state while one partition remains far behind. Conversely, the lag can fall while a move is still finishing. Confirm that the partition is served by a healthy leader, the consumer is processing records, commits are advancing, and the lag slope is moving toward the application's freshness target.
Keep the incident timeline. Note whether the interruption came from leader discovery, member churn, handler slowdown, or movement pressure. That classification determines the next change: improve client shutdown and poll behavior, adjust the movement plan, isolate recovery traffic, or change the storage and compute model being evaluated.
5Where a different storage model changes the decision
The consumer-side diagnosis remains valid on a Kafka-compatible platform. Topics, partitions, consumer groups, offsets, fetches, and commits still describe the application-facing contract. A slow downstream service remains a slow downstream service after a platform migration.
The architectural question appears when repeated broker changes are expensive because durable partition data is tied to each broker. In a traditional shared-nothing layout, adding or replacing brokers can require replica data to be copied before the target placement is useful. That copy competes with production traffic and makes the maintenance window depend on data volume, bandwidth, and recovery state. Consumer groups do not cause that cost, but their freshness SLOs make the cost visible when fetch latency or recovery work delays processing.
A shared-storage, stateless-broker design changes the broker lane. Durable data is stored in shared object storage, while brokers handle protocol processing, routing, and caching. A reassignment can then focus on ownership, leadership, and traffic placement instead of copying the full durable history between local disks. The consumer still needs the same group and lag checks; the difference is that the infrastructure operation may create a shorter and more predictable disruption window.
AutoMQ is a Kafka-compatible cloud-native streaming system built around this storage-compute separation. Its documentation describes stateless brokers and shared storage, but any evaluation should use the workload, WAL type, client behavior, and recovery objectives of the target deployment. Do not substitute an architecture claim for a test: measure the consumer's fetch latency, assignment stability, and lag recovery under the broker changes you actually perform.
The important comparison is therefore operational. If a Redpanda maintenance event produces a short leader transition and the group remains stable, the consumer path may already be healthy. If broker replacement repeatedly creates long data-movement windows that threaten freshness, evaluate whether a platform with stateless brokers removes that specific constraint. Keep the consumer evidence in both cases; it is how you distinguish an architectural limit from a client or application fault.
6A decision tree for the next alert
When lag rises during a Redpanda partition move or broker change, ask these questions in order:
- Did the consumer group membership or assignment change? If yes, inspect the rebalance timeline, member lifecycle, and client deadlines.
- Did the affected partition change leader or report fetch errors? If yes, inspect metadata refresh, leader health, and request latency.
- Did handler or downstream latency change without broker request latency changing? If yes, fix the completion path before changing cluster capacity.
- Did movement work raise broker or network pressure on the same partitions? If yes, review the movement plan, throttles, and recovery controls.
- Is the backlog draining? If yes, estimate recovery against the freshness objective before adding another change during the incident.
This order keeps the investigation tied to causality. It also protects the group from the two most common emergency mistakes: adding members to a group that cannot stay stable, and resetting offsets to hide a lag graph.
7FAQ
7.1Does partition reassignment always trigger a Kafka consumer-group rebalance?
No. Partition reassignment changes broker replicas or leadership. A consumer-group rebalance is driven by group membership, subscription, or assignment events. The two can overlap during maintenance, but verify each timeline independently.
7.2Why can lag increase when the group is stable?
A stable group can still fetch more slowly, process records more slowly, or commit less often. Leader transitions, broker resource pressure, downstream latency, and a hot partition can all increase lag without changing the assignment.
7.3Should I add consumers when Redpanda lag rises during broker scaling?
Only after confirming that the group is stable, the topic has enough partitions to use more members, and downstream capacity can absorb the additional work. If the limiting event is a leader transition or broker movement, more consumers do not remove it.
7.4Which Redpanda command should I use first?
Use the command set supported by your deployed version. rpk group describe helps inspect group state and offsets; partition-move and move-status commands help inspect broker-side movement. Pair those outputs with client logs and application metrics instead of treating a single CLI snapshot as a diagnosis.
7.5What should a migration test measure?
Measure the full work path: assignment stability, leader and fetch behavior, handler time, commit progress, partition-level lag, and recovery time after a broker change. A test that records only reassignment completion misses the consumer freshness objective.
A red lag line during broker maintenance is a prompt to reconstruct the timelines, not a verdict about the consumer group. Start with the partition and the member, separate assignment from ownership, and wait for both infrastructure and application progress to recover before calling the incident closed. If you want to test how a Kafka-compatible platform behaves when broker storage is decoupled from compute, try the AutoMQ open-source project with the same consumer workload and broker-change scenarios you already use.
