fix(subagent): decrypt encrypted model profiles - #4413
Conversation
|
🚦 CI is currently failing on this PR's latest commit. Please fix the failing checks before OpenHands reviews it - this is re-checked automatically once you push a new commit. (A maintainer can also request This is an automated check - no AI was used to generate this comment. |
enyst
left a comment
There was a problem hiding this comment.
Thank you for the PR! 🤔 TBH seeing how we need to keep treading it around only for the profile store (which is the only place who needs it) makes me think that maybe we could have set it in the profile store in the first place...
This is not a suggestion for change right now, though, because we also have a better secrets design issue, and those changes will clean up this... I think. @OpenHands can you find and link that issue? Then tell us WDYT.
|
I am OpenHands (OpenHands-Deepseek-Flash), an AI agent acting on behalf of @enyst. (Note: this is posted through enyst's GitHub token, so it appears under their name — the words here are from the AI agent, not enyst.) Regarding the "better secrets design issue" you asked about — it's #4288 (Design: reference-only credentials and safe runtime delivery): #4288 That's the consolidated design doc enyst contributed to (stdio-MCP, egress-proxy, OS-keyring, iron-proxy), and it's the one whose §2/§9 covers exactly this coupling: the cipher/ On the WDYT: I partly agree the threading is a smell — So: agree it's temporary, agree it's the profile store's job to own at-rest handling, and #4288 is the right place to track killing the whole pattern — link added above. WDYT, good to proceed with this as-is? |
|
Thanks for the review! I agree that threading the cipher through solely for the profile store is not ideal. Given #4288 and the bot’s recommendation, I’m happy to keep this PR focused as-is. Please let me know if you’d like me to make any changes before approval. |
|
🚦 CI is currently failing on this PR's latest commit. Please fix the failing checks before OpenHands reviews it - this is re-checked automatically once you push a new commit. (A maintainer can also request This is an automated check - no AI was used to generate this comment. |
|
Hello, just found this bug as I'm also experiencing it when using subagents:
Is the fix getting merged? Thanks! |
5103299 to
d87a6e8
Compare
|
Thanks @CarlosLanderas for confirming this in real use. @enyst I rebased the PR onto the latest |
|
🤖 OpenHands is reviewing this PR. Head commit: This comment was posted by an AI agent (OpenHands). |
all-hands-bot
left a comment
There was a problem hiding this comment.
This review was created by an AI agent (OpenHands) on behalf of the repository maintainers.
Overall verdict: No material findings. Approve from a review standpoint.
The change correctly threads the conversation cipher through the sub-agent registration path so that encrypted LLM profiles are decrypted before reaching the LLM provider. All four modified production call sites are covered:
agent_definition_to_factorynow acceptscipherand forwards it tostore.load(profile_name, cipher=cipher)— the only place where the cipher is actually needed.register_file_agents/register_plugin_agentsforwardcipherto the factory.LocalConversationpassesself._cipherat all three registration call sites (2 plugin-agent, 1 file-agent).ConversationService._register_agent_definitionsreceives and forwardsself.cipherat both the new-conversation and resume-conversation call sites.
Design note (not blocking): The per-load cipher approach (passing it to store.load rather than setting it on the LLMProfileStore instance) is the right call given that _get_profile_store is a module-level lru_cache singleton shared across conversations. If the cipher were stored on the shared store instance, concurrent conversations with different secret keys but the same profile_store_dir would race on the cipher. The per-load parameter avoids that cleanly. This aligns with enyst's comment that a broader secrets redesign may simplify this later, but the current approach is correct for the existing architecture.
Coverage: register_builtins_agents in preset/default.py still calls agent_definition_to_factory(agent_def) without cipher. This is fine — all built-in agents use model: inherit, so the store.load path where cipher is consumed is never reached for builtins. No change needed.
Test: The regression test (test_agent_definition_to_factory_decrypts_model_profile) exercises the real encrypt→decrypt round trip (no mocks) and verifies the decrypted API key matches the original plaintext. I confirmed it passes.
Risk: LOW. Backward compatible (cipher defaults to None), focused scope, and the fix improves credential handling rather than introducing new surface area.
|
Hey @Shimada666 Fix looks right... A few gaps I'd like to see closed: Sub-conversations don't get a cipher. Two call sites still load profiles with no cipher. F Nothing tests the plumbing. Worth one test asserting the cipher reaches the factory through register_file_agents. Structural, non-blocking: the cipher gets captured in a closure stored in the process-global _agent_factories, and register_agent_if_absent no-ops on an existing name. So whichever conversation registers a name first pins the cipher for every conversation after it. ( I think this is fine). One small ask: both failure modes here are silent. No cipher sends the ciphertext as the API key, and a wrong cipher yields api_key=None. Since FERNET_TOKEN_PREFIX already identifies the value, raising with the profile name would save someone a long debugging session. |
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: openhands <openhands@all-hands.dev>
39fc1bb to
ad61ede
Compare
|
Thanks for the detailed review, @vasco. I've addressed the issues you identified:
I also rebased onto the latest I left the process-global registry lifetime unchanged, as you marked that concern non-blocking; fixing it cleanly would require a broader registry-lifecycle change. Could you please take another look? |
|
Thanks for the quick turnaround @Shimada666. The propagation looks right. Two things mean the new guard doesn't hold yet though.
It also only checks Separately, narrowing
|
Co-authored-by: openhands <openhands@all-hands.dev>
|
Thanks again, @VascoSch92. I addressed the follow-up in
I added regression coverage for case-variant paths, all four secret fields with missing/wrong ciphers, deleted provider connections, and builtin inheritance. 152 relevant tests and the full pre-commit suite pass locally. All completed CI jobs are green; four jobs are still running. Could you please take another look? |
|
🤖 OpenHands is reviewing this PR. Head commit: This comment was posted by an AI agent (OpenHands). |
all-hands-bot
left a comment
There was a problem hiding this comment.
This review was created by an AI agent (OpenHands) on behalf of the repository maintainers.
Overall verdict: No material findings. Approve from a review standpoint.
The change threads the conversation cipher through the sub-agent registration path so encrypted-at-rest LLM profiles are decrypted before their API keys reach a provider. The scope is correct and the implementation is sound:
agent_definition_to_factorygains a keyword-onlycipherand forwards it tostore._load_for_execution(...)— the single point where the profile is actually resolved for execution.register_file_agents/register_plugin_agentsforwardcipher, andLocalConversationpassesself._cipherat all registration plusfork/_bind_conversation_contextsites.ConversationService._register_agent_definitionsreceives and forwards the serverself.cipheron both the create and resume paths.
Correctness of _load_for_execution: The post-load check is the right shape. After load(name, cipher=cipher), it dumps with expose_secrets=True and compares against the on-disk token prefix (gAAAAA). It correctly distinguishes the three cases: (1) loaded value is still a Fernet token → no/absent cipher, reject; (2) stored token but loaded value is None → wrong cipher (decrypt failed in validate_secret), reject; (3) plaintext stored/loaded → accept. A profile saved without encryption still loads fine under a conversation that has a cipher, and vice versa, so plaintext/legacy profiles are not regressed.
Fallback path: FallbackStrategy._bind_cipher + _load_for_execution is a clean minimal addition. The explicit except ProfileDecryptionError: raise is correctly ordered ahead of the broader except (FileNotFoundError, ValueError) (since ProfileDecryptionError subclasses ValueError), so decryption failures surface instead of being silently skipped. Re-raising on an undecryptable fallback profile (rather than sending the raw token to the provider, which is the prior behavior) is the right call.
Design note (non-blocking): storing the cipher per-load (via the factory closure / binding) rather than on the shared _get_profile_store lru_cache singleton is correct — a shared-instance cipher would race across conversations with different secret keys. This matches enyst's earlier observation that a broader secrets redesign can simplify this later.
Coverage: The new regression tests exercise the real encrypt→decrypt round-trip with no mocks and assert the decrypted key equals the original plaintext, including the wrong-cipher/None rejection path. register_builtins_agents still calls agent_definition_to_factory(agent_def) without cipher, which is fine because builtins use model: inherit and never reach the profile-load branch.
Risk: LOW. Keyword-only cipher defaults to None, so all existing call sites remain source-compatible; no version bumps; focused, backwards-compatible scope.
all-hands-bot
left a comment
There was a problem hiding this comment.
This review was created by an AI agent (OpenHands) on behalf of the repository maintainers.
Verdict: No material findings.
This is a focused bug fix: it threads the conversation Cipher through the sub-agent registration paths and rejects profiles whose secrets could not be decrypted, so an encrypted-at-rest API key can no longer silently reach the LLM provider as ciphertext.
What I verified against the workspace:
LLMProfileStore._load_for_executioncorrectly distinguishes the three cases for eachLLM_SECRET_FIELDSentry: (a) plaintext loaded fine, (b) ciphertext still present in the loaded value (no/wrong cipher) →ProfileDecryptionError, (c) stored ciphertext with aNoneloaded value (decrypt failed) →ProfileDecryptionError. The error message leaks neither the secret value nor the key material.Cipheris bound as aPrivateAttr(FallbackStrategy._cipher,LocalConversation._cipher), so it is excluded from model serialization and cannot leak into persisted events/state.- All registration call sites are covered:
agent_definition_to_factory(file/plugin + agent-server_register_agent_definitionsfor both new and resumed conversations),LocalConversation.fork, the delegate executor, and the task manager’s create/resume paths. The remainingregister_builtins_agentscall withoutcipheris fine because built-ins usemodel: inheritand never reach the profile-load branch. fallback_strategy._iter_fallbacksre-raisesProfileDecryptionErrorbefore the broader(FileNotFoundError, ValueError)catch, so a profile with a danglingprovider_connection_idis still skipped while a genuinely undecryptable secret fails loudly.
I ran the touched test suites locally: tests/sdk/llm/test_llm_fallback.py, tests/sdk/llm/test_llm_profile_store.py, and tests/sdk/subagent/test_subagent_registry.py — 142 passed. The regression tests exercise the real encrypt→decrypt round trip (no mocks) and assert on the resulting plaintext.
Risk: LOW. Backward compatible (cipher defaults to None), focused scope, and the change tightens credential handling rather than widening surface area. Not in the eval-risk category (no prompt/agent-behavior change).
|
This comment was posted by an AI agent (OpenHands). |
enyst
left a comment
There was a problem hiding this comment.
Thank you both!
I wonder, sorry, if we really need to do this anymore, since #4492 has introduced Provider Connections, and my hope is that we can remove, not add, but remove most use of cipher in LLM Profiles, because the API keys are in Provider Connections (or will be soon)
We also talked here about the Secrets work, which too, should remove completely I hope, the need for cipher in this codebase. The reason why I care is exactly what you see: with the cipher, it needs to be threaded around in many places, each of those also carry a security risk, and a better design would be safer and less error-prone.
That all said, it’s not the fault of your PR! cc @VascoSch92
The implementation looks good to me, under the older circumstances. Just a heads up, sorry: these fixes will probably not stay long in the code, but for now I’m for merging 😅
HUMAN:
This fixes sub-agent authentication failures when named LLM profiles are encrypted at rest.
AGENT:
Why
When
OH_SECRET_KEYis configured, saved LLM profile secrets are encrypted at rest. File-based sub-agents loaded named profiles without the conversation cipher, so the encrypted API key reached the LLM provider and authentication failed.Summary
Issue Number
Related to #4288.
Fixes #4558
How to Test
uv run pytest -q tests/sdk/subagent/test_subagent_registry.py tests/sdk/conversation/test_local_conversation_plugins.py— 81 passed.uv run pytest -q tests/agent_server/test_conversation_service.py tests/agent_server/test_conversation_service_plugin.py— 109 passed.include_secrets=Trueand aCipher, load anevidence-coderagent from an explicit plugin throughLocalConversation, and instantiate the registered factory. The smoke test returned modelopenai/gpt-5.6-lunaand confirmed the API key matched the original plaintext.Video/Screenshots
Not applicable; this is an SDK credential-loading fix. The deterministic regression test and plugin smoke-test results are described above.
Type
Notes
Documentation: OpenHands/docs#696