Is your feature request related to a problem?
Environment
- OSO Kafka backup version: v0.22.0
- Kafka: Strimzi-managed Kafka 3.8.0
- Target cluster: 4 brokers
- Restore type: Bulk disaster-recovery restore
- Concurrent partitions: 8
- Approximate records: 1.35 billion
- Checkpointing: Enabled
Summary
The Kafka backup/restore tool contains an internal circuit breaker in kafka-backup-core. Its configuration includes:
failure_threshold
reset_timeout
success_threshold
These values are currently hardcoded and cannot be configured through YAML, CLI flags, or environment variables.
Problem
During a large Kafka restore with version v0.22.0, we repeatedly observed:
Component kafka became Degraded
Component kafka recovered
Packet captures showed approximately 45-60 seconds of complete bidirectional silence across all broker connections during these events. Kafka broker logs remained clean, indicating that the pause occurred inside the client rather than in Kafka or the network.
This behavior is consistent with the global Kafka circuit breaker entering the Open state and waiting for its reset timeout.
Because the breaker appears to cover the entire Kafka operation, a transient failure affecting one partition can pause all concurrently restored partitions.
Our restore contained approximately 1.35 billion records and used:
restore:
max_concurrent_partitions: 8
produce_batch_size: 500
rate_limit_records_per_sec: 100000
checkpoint_state: /backup/restore-checkpoint.json
The repeated circuit-breaker pauses significantly increased the total restore duration.
Related work
PR #34, released as v0.8.2, exposed several previously hardcoded networking and resilience settings, including:
connections_per_broker
produce_acks
produce_timeout_ms
Connection retry handling
Socket read/write timeouts
Circuit-breaker behavior appears to be a related area that is still not configurable.
Describe the solution you'd like
Requested feature
Please expose circuit-breaker settings in the configuration. For example:
restore:
circuit_breaker:
enabled: true
failure_threshold: 20
reset_timeout_ms: 2000
success_threshold: 1
At minimum, please provide an option to disable the circuit breaker for restore operations:
restore:
circuit_breaker:
enabled: false
Restore operations already support checkpointing and retry behavior. For large, resumable, one-shot restores, operators may reasonably prefer to tolerate isolated transient failures without globally pausing the entire restore.
A further improvement would be to scope circuit breakers per broker or partition instead of globally. A failure affecting one partition should not suspend unrelated restore work.
Suggested configuration
For a bulk restore, reasonable starting values might be:
failure_threshold: 15
reset_timeout_ms: 2000
success_threshold: 1
Existing production defaults could remain unchanged when no override is provided.
Describe alternatives you've considered
No response
Use case
No response
Additional context
No response
Is your feature request related to a problem?
Environment
Summary
The Kafka backup/restore tool contains an internal circuit breaker in
kafka-backup-core. Its configuration includes:These values are currently hardcoded and cannot be configured through YAML, CLI flags, or environment variables.
Problem
During a large Kafka restore with version
v0.22.0, we repeatedly observed:Packet captures showed approximately 45-60 seconds of complete bidirectional silence across all broker connections during these events. Kafka broker logs remained clean, indicating that the pause occurred inside the client rather than in Kafka or the network.
This behavior is consistent with the global Kafka circuit breaker entering the Open state and waiting for its reset timeout.
Because the breaker appears to cover the entire Kafka operation, a transient failure affecting one partition can pause all concurrently restored partitions.
Our restore contained approximately 1.35 billion records and used:
The repeated circuit-breaker pauses significantly increased the total restore duration.
Related work
PR #34, released as
v0.8.2, exposed several previously hardcoded networking and resilience settings, including:connections_per_brokerproduce_acksproduce_timeout_msConnection retry handlingSocket read/write timeoutsCircuit-breaker behavior appears to be a related area that is still not configurable.
Describe the solution you'd like
Requested feature
Please expose circuit-breaker settings in the configuration. For example:
At minimum, please provide an option to disable the circuit breaker for restore operations:
Restore operations already support checkpointing and retry behavior. For large, resumable, one-shot restores, operators may reasonably prefer to tolerate isolated transient failures without globally pausing the entire restore.
A further improvement would be to scope circuit breakers per broker or partition instead of globally. A failure affecting one partition should not suspend unrelated restore work.
Suggested configuration
For a bulk restore, reasonable starting values might be:
Existing production defaults could remain unchanged when no override is provided.
Describe alternatives you've considered
No response
Use case
No response
Additional context
No response