Five failure modes I hit building a guide-driven migration tool — are these known? #2410
LucasAbud11
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I spent the last six weeks building a CLI that reads a published
migration guide and finds what breaks in a Python codebase
(github.com/LucasAbud11/api-drift). I'm 18 and did this alone, so I
expect some of what follows is well-known to people who've been at this
longer — I'd genuinely like to know which.
I ended up at roughly the architecture codemod describes: deterministic
search for candidates, model only for judgment. Five things surprised me,
and I can't find them written up:
Fixes that are individually correct and collectively insufficient.**
On a real repo, mine renamed three sites correctly and would have shipped
code that crashes at startup, because a fourth change 200 lines away was
required. Both verification tiers passed. I now detect the coupling and
decline the whole group, but only after the failure taught me to.
Verification proves self-consistency, not sufficiency.** A fix can
parse, match its claimed original line, and resolve its import, and still
be wrong. Nothing static catches "this change was necessary but not
enough."
Fix generation gets refused based on repository content.** Returns
stop_reason=refusal reproducibly on security tooling — two independent
repos, one of them defensive (an MCP misconfiguration scanner). Nothing
to do with the migration itself.
Coverage metrics wrong in both directions.** My guard understated
gaps through three separate bugs, then turned out to also overstate
coverage on facts whose only match was coincidental. Took four
corrections before the number meant anything.
Deprecation warnings eat much of the use case.** Where a library
deprecates in place, the warning names the file, line, and replacement —
better than any pattern I could derive. Where it doesn't, there's often
no identifier to grep for either. The addressable middle is narrower than
I assumed going in.
Best result so far: 405-file codebase, Pydantic v1→v2, 174 verified
fixes and 55 flagged for human review — including two of its own
proposals caught claiming original text that wasn't in the file.
Are 1-4 things you've already solved, or fundamental to the approach?
And is 5 how you'd characterize the market shape too?
All reactions