Skip to content

Calibration: concurrent sync_read from get_status() and recording loop causes "[TxRxResult] Incorrect status packet!" #121

Description

@matsujirushi

Description

During calibration (range-of-motion recording) of an SO-101 Follower, reading positions intermittently fails with:

WARNING 2026-09-29 10:21:48 alibrate.py:167 Failed to read positions: Failed to sync read 'Present_Position' on ids=[1, 2, 3, 4, 5, 6] after 1 tries. [TxRxResult] Incorrect status packet!
WARNING 2026-09-29 10:21:48 alibrate.py:461 Error reading positions during recording: Failed to sync read 'Present_Position' on ids=[1, 2, 3, 4, 5, 6] after 1 tries. [TxRxResult] Incorrect status packet!

Both warnings are logged at the same time, one from get_status() and one from the recording loop. The servos and the wiring appear to be fine. The problem looks like a host-side race: two threads use the same serial port at the same time.

Root cause (analysis)

In lelab/calibrate.py, two code paths call bus.sync_read("Present_Position") on the same MotorsBus:

  1. Recording loop: reads positions, then runs time.sleep(0.05) (about 20 Hz). On error it logs Error reading positions during recording and runs time.sleep(0.2).
  2. get_status(): while status is "recording", it runs its own sync_read. The frontend calls it through /calibration-status, which Calibration.tsx polls every 200 ms (setInterval(pollStatus, 200)). FastAPI runs this handler in a different thread from the recording loop.

The feetech SDK (scservo_sdk) is not thread-safe:

  • Its only guard is the port.is_using flag, and the check-then-set on that flag is not atomic.
  • GroupSyncRead.rxPacket() calls readRx() once per ID. Each rxPacket() resets is_using = False after every single-servo reply. Another thread can therefore start a transmission while a sync read is still receiving the replies from IDs 2-6.
  • txPacket() calls port.clearPort(), which flushes the RX buffer. This discards bytes that the other thread is still waiting for, and its parser then returns COMM_RX_CORRUPT ("Incorrect status packet").

The retry logic only handles Port is in use, so this error is re-raised immediately.

Evidence (logic analyzer capture)

I captured the half-duplex bus with a Saleae Logic Pro 16 at 1 Mbps (8N1) during calibration:

  • 154 SYNC_READ requests (FF FF FE 0A 82 38 02 01 02 03 04 05 06 26) and 919 status packets were decoded.
  • All checksums were valid, the error byte was 0x00 in every status packet, and there were no framing errors. Supply voltage was stable (~4.95-4.97 V at the probe point), and the analog edges were clean.
  • At the time of the failure:
    • t=5.665198 s: SYNC_READ, and all 6 servos replied correctly (finished at ~5.666000 s).
    • t=5.666492 s: a second, identical SYNC_READ only 1.3 ms later, and all 6 servos again replied correctly.
    • Then no traffic for 196.5 ms, which matches the time.sleep(0.2) error back-off.
  • Request intervals: many ~52 ms (the recording loop) plus irregular 10-45 ms intervals. These are consistent with the 200 ms status polling being interleaved with the loop.

So the servos sent valid replies to both requests, and both reads still failed on the host. This points to concurrent access to the port rather than a problem with the servos or the wiring.

Environment

  • lerobot: v0.6.0 (as pinned in pyproject)
  • feetech-servo-sdk: 1.0.0
  • Robot: SO-101 Follower (STS3215 x6)

Steps to reproduce

  1. Start LeLab and open the Calibration page for an SO-101 Follower.
  2. Start calibration and move the joints during range-of-motion recording.
  3. The warnings above appear intermittently.

Suggested fix

Option A (preferred): make the recording loop the only reader of the bus. It would store the latest positions in a shared variable, and get_status() would return those cached positions instead of calling sync_read.

Option B: protect every access to the bus with a threading.Lock, shared by the recording loop, get_status(), and any other calibration step.

# sketch of option A
with self._state_lock:
    self._latest_positions = positions   # in recording loop

# in get_status()
with self._state_lock:
    positions = dict(self._latest_positions)

Related: huggingface/lerobot#2793 (similar symptom caused by concurrent bus access), huggingface/lerobot#1010 (same message, but a hardware cause), #72, #84.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions