Skip to content

Expose circuit breaker configuration for restore workloads #197

Description

@obacak

Is your feature request related to a problem?

Environment

  • OSO Kafka backup version: v0.22.0
  • Kafka: Strimzi-managed Kafka 3.8.0
  • Target cluster: 4 brokers
  • Restore type: Bulk disaster-recovery restore
  • Concurrent partitions: 8
  • Approximate records: 1.35 billion
  • Checkpointing: Enabled

Summary

The Kafka backup/restore tool contains an internal circuit breaker in kafka-backup-core. Its configuration includes:

failure_threshold
reset_timeout
success_threshold

These values are currently hardcoded and cannot be configured through YAML, CLI flags, or environment variables.

Problem

During a large Kafka restore with version v0.22.0, we repeatedly observed:

Component kafka became Degraded
Component kafka recovered

Packet captures showed approximately 45-60 seconds of complete bidirectional silence across all broker connections during these events. Kafka broker logs remained clean, indicating that the pause occurred inside the client rather than in Kafka or the network.

This behavior is consistent with the global Kafka circuit breaker entering the Open state and waiting for its reset timeout.

Because the breaker appears to cover the entire Kafka operation, a transient failure affecting one partition can pause all concurrently restored partitions.

Our restore contained approximately 1.35 billion records and used:

restore:
  max_concurrent_partitions: 8
  produce_batch_size: 500
  rate_limit_records_per_sec: 100000
  checkpoint_state: /backup/restore-checkpoint.json

The repeated circuit-breaker pauses significantly increased the total restore duration.

Related work

PR #34, released as v0.8.2, exposed several previously hardcoded networking and resilience settings, including:

  • connections_per_broker
  • produce_acks
  • produce_timeout_ms
  • Connection retry handling
  • Socket read/write timeouts

Circuit-breaker behavior appears to be a related area that is still not configurable.

Describe the solution you'd like

Requested feature

Please expose circuit-breaker settings in the configuration. For example:

restore:
  circuit_breaker:
    enabled: true
    failure_threshold: 20
    reset_timeout_ms: 2000
    success_threshold: 1

At minimum, please provide an option to disable the circuit breaker for restore operations:

restore:
  circuit_breaker:
    enabled: false

Restore operations already support checkpointing and retry behavior. For large, resumable, one-shot restores, operators may reasonably prefer to tolerate isolated transient failures without globally pausing the entire restore.

A further improvement would be to scope circuit breakers per broker or partition instead of globally. A failure affecting one partition should not suspend unrelated restore work.

Suggested configuration

For a bulk restore, reasonable starting values might be:

failure_threshold: 15
reset_timeout_ms: 2000
success_threshold: 1

Existing production defaults could remain unchanged when no override is provided.

Describe alternatives you've considered

No response

Use case

No response

Additional context

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions