Skip to content

[Bug]: Prompt-based tool calling fails on Bedrock OpenAI models (gpt-oss, GPT-5.6, GPT-6 Sol/Luna/Astra) because stop words are sent #5290

Description

@kimnamu

Thanks for keeping model quirks in one place in model_features.py; the o1/o3/grok stop-word entries made this quick to trace.

Is there an existing issue for the same bug?

  • I have searched existing issues and this is not a duplicate.

Searched issues and PRs: no match.

Bug Description

With native_tool_calling=False, chat-completions calls with tools to OpenAI models on Bedrock (gpt-oss-120b, GPT-5.6 Sol, GPT-6 Sol/Luna/Astra) fail with a 400. The SDK adds stop=["</function"] (mixins/non_native_fc.py:59-61), LiteLLM sends it as Converse stopSequences, and these models reject that field. Claude on Bedrock accepts the stop word. gpt-oss hits this by default, GPT-5.6/GPT-6 with api_mode="chat" (auto sends them to the Responses path, which adds no stop words).

Expected Behavior

Prompt-based tool calling works, as it does for o1/o3/grok-4 via SUPPORTS_STOP_WORDS_FALSE_MODELS.

Actual Behavior

uv run python repro.py (code sample below) on main @ 5b36cac:

ERR bedrock/openai.gpt-oss-120b-1:0 BedrockException - {"message":"This model doesn't support the stopSequences field. Remove stopSequences and try again."}
ERR bedrock/us.openai.gpt-5.6-sol BedrockException - {"message":"This model doesn't support the stopSequences field. Remove stopSequences and try again."}
ERR bedrock/us.openai.gpt-6-sol BedrockException - {"message":"This model doesn't support the stopSequences field. Remove stopSequences and try again."}
ERR bedrock/us.openai.gpt-6-luna BedrockException - {"message":"This model doesn't support the stopSequences field. Remove stopSequences and try again."}
ERR bedrock/us.openai.gpt-6-astra BedrockException - {"message":"This model doesn't support the stopSequences field. Remove stopSequences and try again."}
OK  bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 tool_calls: ['think']

Steps to Reproduce

  1. uv sync --extra boto3, Bedrock credentials in the environment.
  2. uv run python repro.py.

Acceptance Criteria

  • get_features(...).supports_stop_words is False for bedrock/openai.gpt-oss-120b-1:0, bedrock/us.openai.gpt-5.6-sol and bedrock/{us,global}.openai.gpt-6-{sol,luna,astra}
  • The repro no longer returns the stopSequences 400
  • Claude on Bedrock still gets the stop word

Installation Method

uv sync on main (LiteLLM 1.93.0 from uv.lock)

SDK Version

main @ 5b36cac (openhands-sdk 1.49.5)

Python Version

3.13.14

Model Name (if applicable)

bedrock/openai.gpt-oss-120b-1:0, bedrock/us.openai.gpt-5.6-sol, bedrock/us.openai.gpt-6-sol, bedrock/us.openai.gpt-6-luna, bedrock/us.openai.gpt-6-astra

Operating System

MacOS

Minimal Code Sample

from openhands.sdk import LLM, Message, TextContent
from openhands.sdk.tool.builtins.think import ThinkTool

# reasoning_effort=None: the locked LiteLLM 1.93.0 would otherwise turn the
# default "high" into a `thinking` field that GPT-5.6/GPT-6 reject.
MODELS = {
    "bedrock/openai.gpt-oss-120b-1:0": {},
    "bedrock/us.openai.gpt-5.6-sol": {"reasoning_effort": None},
    "bedrock/us.openai.gpt-6-sol": {"reasoning_effort": None},
    "bedrock/us.openai.gpt-6-luna": {"reasoning_effort": None},
    "bedrock/us.openai.gpt-6-astra": {"reasoning_effort": None},
    "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0": {},
}
msgs = [
    Message(role="system", content=[TextContent(text="You are an agent. Use the think tool once, then answer.")]),
    Message(role="user", content=[TextContent(text="Call the think tool with thought='2+2', then tell me the result.")]),
]
for model, extra in MODELS.items():
    llm = LLM(model=model, usage_id="repro", api_mode="chat", native_tool_calling=False,
              aws_region_name="us-east-1", num_retries=0, **extra)
    try:
        r = llm.completion(msgs, tools=ThinkTool.create())
        print("OK ", model, "tool_calls:", [tc.name for tc in r.message.tool_calls or []])
    except Exception as e:
        print("ERR", model, str(e)[-120:])

Screenshots and Additional Context

Now Proposed
Bedrock openai.gpt-*, prompt-based tools 400 stopSequences no stop word sent, call succeeds
Claude on Bedrock, prompt-based tools stop word sent ✅ unchanged ✅
native_tool_calling=True (default) no stop word ✅ unchanged ✅

Workaround today: disable_stop_word=True.

Proposed change and notes

One entry in SUPPORTS_STOP_WORDS_FALSE_MODELS, like #147 did for grok-code-fast-1:

     # DeepSeek R1 family
     "deepseek-r1-0528",
+    # OpenAI models on Bedrock (gpt-oss, gpt-5.x, gpt-6) reject stopSequences
+    "openai.gpt-",
 ]

The substring only matches the dotted id form, so openai/gpt-5 style ids are unchanged. It also matches bedrock_mantle/openai.gpt-* and LiteLLM's three oci/openai.gpt-5* ids, which I did not run through the SDK.

With the SDK default reasoning_effort="high", the locked LiteLLM 1.93.0 adds an Anthropic-style thinking field and GPT-5.6/GPT-6 answer 400 Unknown parameter: 'thinking'. That is a separate problem (#5292), so the repro sets reasoning_effort=None for those rows.

I used Claude Code to trace the code path, search the tracker and run the repro. I'll open a draft PR with this change and tests.


OpenHands AI triage

The following comments and acceptance criteria were added by the OpenHands AI agent.

Triage

Confirmed reproducible bug, not a duplicate (the only other stopSequences hits are this issue and its fix PR #5291). Root cause is in the repository's authoritative model-capability registry: SUPPORTS_STOP_WORDS_FALSE_MODELS in openhands-sdk/openhands/sdk/llm/utils/model_features.py feeds get_features(...).supports_stop_words, which mixins/non_native_fc.py consults before adding stop=["</function"]. model_matches is a case-insensitive substring test on the full raw model id, so the dotted-id token used by Bedrock-style routes is the right granularity; the issue's openai.gpt- proposal matches the precedent set for grok-code-fast-1 (#147) and the substring convention established in #844.

Scope: prompt-based tool calling (native_tool_calling=False) against OpenAI models reached through Bedrock dotted ids (openai.gpt-*, {us,global}.openai.gpt-*). Fix belongs in the feature registry, not as a special case in llm.py.

Accepted side effect, bounded: the same dotted-id token necessarily also matches bedrock_mantle/openai.gpt-* and oci/openai.gpt-5*. Those routes were not exercised live by the reporter, and verifying them is explicitly out of scope here; the change is accepted because it only suppresses an optional prompt-format stop word, and non-dotted ids (e.g. openai/gpt-5) are untouched. Reviewers should treat a narrower, Bedrock-only enumeration as acceptable if it covers every id below.

Explicit non-goals: the separate reasoning_effort="high" → Anthropic-style thinking 400 on GPT-5.6/GPT-6 (the issue calls this out as its own problem); any change to llm.py, the public LLM field set, or native tool calling / Responses-path behavior.

Acceptance Criteria

  • get_features(...).supports_stop_words is False for bedrock/openai.gpt-oss-120b-1:0, bedrock/us.openai.gpt-5.6-sol, bedrock/us.openai.gpt-6-sol, bedrock/us.openai.gpt-6-luna, bedrock/us.openai.gpt-6-astra, and the global. inference-profile variants bedrock/global.openai.gpt-6-sol, bedrock/global.openai.gpt-6-luna, bedrock/global.openai.gpt-6-astra.
  • The change is expressed in the existing capability registry (SUPPORTS_STOP_WORDS_FALSE_MODELS / get_features), not as a branch in llm.py, and no public LLM field or default changes.
  • Non-dotted OpenAI ids are unchanged and still report supports_stop_words is True: openai/gpt-5, openai/gpt-oss-120b, litellm_proxy/openai/gpt-5, gpt-4o, unknown-model.
  • Bedrock Claude is unchanged and still reports supports_stop_words is True, so prompt-based tool calling to bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 still sends the stop word.
  • Existing opt-outs still hold: o1, o1-2024-12-17, o3, grok-4-0709, grok-code-fast-1, deepseek-r1-0528.
  • Existing escape hatches still win over the new entry: disable_stop_word=True and capability_overrides={"supports_stop_words": True} each keep the stop word for the listed ids.
  • Regression coverage lands in tests/sdk/llm/test_model_features.py (the test_supports_stop_words_false_models parametrization plus a Claude-on-Bedrock keep case) and uv run pytest tests/sdk/llm/test_model_features.py passes.
  • Live validation with the issue's repro (native_tool_calling=False, api_mode="chat"): the five reported Bedrock ids complete without the stopSequences 400, the Converse request body contains no stopSequences, and Claude Haiku on Bedrock still succeeds with the stop word present.
  • Unchanged paths verified: with default native tool calling, and on the Responses path that api_mode="auto" selects for gpt-5.6/gpt-6, no stop word is added before or after the change.
  • No AWS/Bedrock credentials or account identifiers are committed; unit-level criteria are satisfied by get_features assertions that need no live credentials.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:highFor bugs, affecting nearly all users and degrading performance or UX.ready-for-devIssue meets development readiness criteria

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions