Release v1.44.0 - #4628
Conversation
Co-authored-by: openhands <openhands@all-hands.dev>
|
Hi! I started running the integration tests on your PR. You will receive a comment with the results shortly. |
1 similar comment
|
Hi! I started running the integration tests on your PR. You will receive a comment with the results shortly. |
|
Hi! I started running the behavior tests on your PR. You will receive a comment with the results shortly. |
1 similar comment
|
Hi! I started running the behavior tests on your PR. You will receive a comment with the results shortly. |
🔒 Release Security Scan🔒 Approval drift (time-of-check vs time-of-use)❌ 3 finding(s) Baseline:
Audited 11 PR(s): 8 clean, 3 flagged, 0 un-auditable. 📦 Supply-chain dependency diff✅ no findings Baseline: OSV: no known vulns across 6 new/bumped dep(s). Bumped dependencies
Internal `openhands-*` bumps
Deterministic scanners: approval-drift + supply-chain dependency diff. Read-only; no PR code executed. |
2 similar comments
🔒 Release Security Scan🔒 Approval drift (time-of-check vs time-of-use)❌ 3 finding(s) Baseline:
Audited 11 PR(s): 8 clean, 3 flagged, 0 un-auditable. 📦 Supply-chain dependency diff✅ no findings Baseline: OSV: no known vulns across 6 new/bumped dep(s). Bumped dependencies
Internal `openhands-*` bumps
Deterministic scanners: approval-drift + supply-chain dependency diff. Read-only; no PR code executed. |
🔒 Release Security Scan🔒 Approval drift (time-of-check vs time-of-use)❌ 3 finding(s) Baseline:
Audited 11 PR(s): 8 clean, 3 flagged, 0 un-auditable. 📦 Supply-chain dependency diff✅ no findings Baseline: OSV: no known vulns across 6 new/bumped dep(s). Bumped dependencies
Internal `openhands-*` bumps
Deterministic scanners: approval-drift + supply-chain dependency diff. Read-only; no PR code executed. |
Python API breakage checks — ✅ PASSEDResult: ✅ PASSED |
REST API breakage checks (OpenAPI) — ✅ PASSEDResult: ✅ PASSED |
🧪 Integration Tests ResultsOverall Success Rate: 97.7% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_deepseek_deepseek_v4_flash
Skipped Tests:
litellm_proxy_anthropic_claude_sonnet_4_6
Failed Tests:
litellm_proxy_gemini_3.1_pro_preview
litellm_proxy_openai_gpt_5.5
litellm_proxy_minimax_MiniMax_M2.7
Skipped Tests:
|
|
👋 This PR needs a couple of things fixed before OpenHands can review it:
Push an update once this is addressed and this check re-runs automatically. This is an automated check - no AI was used to generate this comment. |
🔄 Running Examples with
|
| Example | Status | Duration | Cost |
|---|---|---|---|
| 01_standalone_sdk/02_custom_tools.py | ✅ PASS | 24.2s | $0.03 |
| 01_standalone_sdk/03_activate_skill.py | ✅ PASS | 22.1s | $0.03 |
| 01_standalone_sdk/05_use_llm_registry.py | ✅ PASS | 11.6s | $0.01 |
| 01_standalone_sdk/07_mcp_integration.py | ✅ PASS | 37.7s | $0.03 |
| 01_standalone_sdk/09_pause_example.py | ✅ PASS | 12.5s | $0.01 |
| 01_standalone_sdk/10_persistence.py | ✅ PASS | 27.1s | $0.02 |
| 01_standalone_sdk/11_async.py | ✅ PASS | 34.1s | $0.04 |
| 01_standalone_sdk/12_custom_secrets.py | ✅ PASS | 17.3s | $0.01 |
| 01_standalone_sdk/13_get_llm_metrics.py | ✅ PASS | 34.8s | $0.05 |
| 01_standalone_sdk/14_context_condenser.py | ✅ PASS | 3m 42s | $0.22 |
| 01_standalone_sdk/17_image_input.py | ✅ PASS | 21.7s | $0.02 |
| 01_standalone_sdk/18_send_message_while_processing.py | ✅ PASS | 23.8s | $0.01 |
| 01_standalone_sdk/19_llm_routing.py | ✅ PASS | 19.2s | $0.02 |
| 01_standalone_sdk/20_stuck_detector.py | ✅ PASS | 17.2s | $0.02 |
| 01_standalone_sdk/21_generate_extraneous_conversation_costs.py | ✅ PASS | 13.4s | $0.00 |
| 01_standalone_sdk/22_anthropic_thinking.py | ✅ PASS | 18.1s | $0.01 |
| 01_standalone_sdk/23_responses_reasoning.py | ✅ PASS | 50.6s | $0.01 |
| 01_standalone_sdk/24_planning_agent_workflow.py | ✅ PASS | 3m 11s | $0.23 |
| 01_standalone_sdk/25_agent_delegation.py | ✅ PASS | 55.0s | $0.06 |
| 01_standalone_sdk/26_custom_visualizer.py | ✅ PASS | 18.9s | $0.02 |
| 01_standalone_sdk/28_ask_agent_example.py | ❌ FAIL Exit code 1 |
19.1s | -- |
| 01_standalone_sdk/29_llm_streaming.py | ✅ PASS | 35.2s | $0.02 |
| 01_standalone_sdk/30_tom_agent.py | ✅ PASS | 14.7s | $0.01 |
| 01_standalone_sdk/31_iterative_refinement.py | ✅ PASS | 3m 25s | $0.22 |
| 01_standalone_sdk/32_configurable_security_policy.py | ✅ PASS | 20.7s | $0.01 |
| 01_standalone_sdk/33_hooks/main.py | ✅ PASS | 38.6s | $0.04 |
| 01_standalone_sdk/34_critic_example.py | ✅ PASS | 6m 14s | $0.42 |
| 01_standalone_sdk/36_event_json_to_openai_messages.py | ✅ PASS | 13.9s | $0.01 |
| 01_standalone_sdk/37_llm_profile_store/main.py | ✅ PASS | 9.2s | $0.00 |
| 01_standalone_sdk/38_browser_session_recording.py | ✅ PASS | 43.0s | $0.04 |
| 01_standalone_sdk/39_llm_fallback.py | ✅ PASS | 18.7s | $0.01 |
| 01_standalone_sdk/40_acp_agent_example.py | ✅ PASS | 48.8s | $0.13 |
| 01_standalone_sdk/41_task_tool_set.py | ✅ PASS | 23.7s | $0.01 |
| 01_standalone_sdk/42_file_based_subagents.py | ✅ PASS | 1m 5s | $0.06 |
| 01_standalone_sdk/44_model_switching_in_convo.py | ✅ PASS | 11.5s | $0.01 |
| 01_standalone_sdk/45_parallel_tool_execution.py | ✅ PASS | 4m 29s | $0.58 |
| 01_standalone_sdk/46_agent_settings.py | ✅ PASS | 12.5s | $0.01 |
| 01_standalone_sdk/47_defense_in_depth_security.py | ✅ PASS | 4.2s | $0.00 |
| 01_standalone_sdk/48_conversation_fork.py | ✅ PASS | 21.8s | $0.00 |
| 01_standalone_sdk/49_switch_llm_tool.py | ✅ PASS | 10.1s | $0.02 |
| 01_standalone_sdk/50_async_cancellation.py | ✅ PASS | 14.4s | $0.01 |
| 01_standalone_sdk/51_agent_hooks/main.py | ✅ PASS | 36.5s | $0.04 |
| 01_standalone_sdk/52_dynamic_workflow.py | ✅ PASS | 3m 37s | $0.11 |
| 01_standalone_sdk/53_client_defined_tools.py | ✅ PASS | 13.3s | $0.01 |
| 01_standalone_sdk/54_goal_completion_loop.py | ✅ PASS | 29.4s | $0.03 |
| 01_standalone_sdk/55_persistent_memory.py | ✅ PASS | 27.0s | $0.02 |
| 01_standalone_sdk/56_structured_output.py | ✅ PASS | 41.5s | $0.06 |
| 01_standalone_sdk/57_prompt_hooks/main.py | ✅ PASS | 10.8s | $0.00 |
| 02_remote_agent_server/01_convo_with_local_agent_server.py | ✅ PASS | 37.7s | $0.02 |
| 02_remote_agent_server/02_convo_with_docker_sandboxed_server.py | ✅ PASS | 1m 34s | $0.04 |
| 02_remote_agent_server/03_browser_use_with_docker_sandboxed_server.py | ✅ PASS | 1m 46s | $0.05 |
| 02_remote_agent_server/04_convo_with_api_sandboxed_server.py | ✅ PASS | 1m 36s | $0.03 |
| 02_remote_agent_server/06_custom_tool/main.py | ✅ PASS | 6m 5s | $0.09 |
| 02_remote_agent_server/07_convo_with_cloud_workspace.py | ✅ PASS | 42.2s | $0.02 |
| 02_remote_agent_server/08_convo_with_apptainer_sandboxed_server.py | ✅ PASS | 4m 7s | $0.03 |
| 02_remote_agent_server/09_acp_agent_with_remote_runtime.py | ✅ PASS | 1m 48s | $0.36 |
| 02_remote_agent_server/10_cloud_workspace_share_credentials.py | ✅ PASS | 38.7s | $0.02 |
| 02_remote_agent_server/11_conversation_fork.py | ✅ PASS | 45.0s | $0.00 |
| 02_remote_agent_server/12_settings_and_secrets_api.py | ✅ PASS | 2m 26s | $0.02 |
| 02_remote_agent_server/13_workspace_get_llm.py | ✅ PASS | 30.5s | $0.01 |
| 02_remote_agent_server/14_client_defined_tools.py | ✅ PASS | 34.7s | $0.02 |
| 02_remote_agent_server/15_openai_compatible_gateway.py | ✅ PASS | 26.9s | $0.01 |
| 02_remote_agent_server/16_deferred_init.py | ✅ PASS | 22.0s | $0.00 |
| 04_llm_specific_tools/01_gpt5_apply_patch_preset.py | ✅ PASS | 28.3s | $0.02 |
| 04_llm_specific_tools/02_gemini_file_tools.py | ✅ PASS | 49.5s | $0.08 |
| 05_skills_and_plugins/01_loading_agentskills/main.py | ✅ PASS | 18.6s | $0.01 |
| 05_skills_and_plugins/02_loading_plugins/main.py | ✅ PASS | 19.4s | $0.03 |
| 05_skills_and_plugins/04_mixed_marketplace_skills/main.py | ✅ PASS | 5.2s | $0.00 |
❌ Some tests failed
Total: 68 | Passed: 67 | Failed: 1 | Total Cost: $3.60
Failed examples:
- examples/01_standalone_sdk/28_ask_agent_example.py: Exit code 1
🔄 Running Examples with
|
| Example | Status | Duration | Cost |
|---|---|---|---|
| 01_standalone_sdk/02_custom_tools.py | ✅ PASS | 27.0s | $0.02 |
| 01_standalone_sdk/03_activate_skill.py | ✅ PASS | 21.7s | $0.01 |
| 01_standalone_sdk/05_use_llm_registry.py | ✅ PASS | 11.3s | $0.00 |
| 01_standalone_sdk/07_mcp_integration.py | ✅ PASS | 38.5s | $0.02 |
| 01_standalone_sdk/09_pause_example.py | ✅ PASS | 14.9s | $0.01 |
| 01_standalone_sdk/10_persistence.py | ✅ PASS | 30.6s | $0.03 |
| 01_standalone_sdk/11_async.py | ✅ PASS | 32.7s | $0.03 |
| 01_standalone_sdk/12_custom_secrets.py | ✅ PASS | 18.6s | $0.01 |
| 01_standalone_sdk/13_get_llm_metrics.py | ✅ PASS | 27.5s | $0.03 |
| 01_standalone_sdk/14_context_condenser.py | ✅ PASS | 2m 40s | $0.16 |
| 01_standalone_sdk/17_image_input.py | ✅ PASS | 26.3s | $0.02 |
| 01_standalone_sdk/18_send_message_while_processing.py | ✅ PASS | 25.0s | $0.01 |
| 01_standalone_sdk/19_llm_routing.py | ✅ PASS | 19.5s | $0.01 |
| 01_standalone_sdk/20_stuck_detector.py | ✅ PASS | 16.6s | $0.01 |
| 01_standalone_sdk/21_generate_extraneous_conversation_costs.py | ✅ PASS | 11.7s | $0.00 |
| 01_standalone_sdk/22_anthropic_thinking.py | ✅ PASS | 14.3s | $0.01 |
| 01_standalone_sdk/23_responses_reasoning.py | ✅ PASS | 1m 13s | $0.01 |
| 01_standalone_sdk/24_planning_agent_workflow.py | ✅ PASS | 4m 21s | $0.33 |
| 01_standalone_sdk/25_agent_delegation.py | ✅ PASS | 52.8s | $0.05 |
| 01_standalone_sdk/26_custom_visualizer.py | ✅ PASS | 23.2s | $0.03 |
| 01_standalone_sdk/28_ask_agent_example.py | ❌ FAIL Exit code 1 |
17.8s | -- |
| 01_standalone_sdk/29_llm_streaming.py | ✅ PASS | 39.5s | $0.02 |
| 01_standalone_sdk/30_tom_agent.py | ✅ PASS | 10.4s | $0.00 |
| 01_standalone_sdk/31_iterative_refinement.py | ✅ PASS | 1m 27s | $0.08 |
| 01_standalone_sdk/32_configurable_security_policy.py | ✅ PASS | 18.6s | $0.02 |
| 01_standalone_sdk/33_hooks/main.py | ✅ PASS | 32.6s | $0.04 |
| 01_standalone_sdk/34_critic_example.py | ✅ PASS | 8m 10s | $0.65 |
| 01_standalone_sdk/36_event_json_to_openai_messages.py | ✅ PASS | 11.6s | $0.00 |
| 01_standalone_sdk/37_llm_profile_store/main.py | ✅ PASS | 6.4s | $0.00 |
| 01_standalone_sdk/38_browser_session_recording.py | ✅ PASS | 33.8s | $0.03 |
| 01_standalone_sdk/39_llm_fallback.py | ✅ PASS | 11.9s | $0.00 |
| 01_standalone_sdk/40_acp_agent_example.py | ✅ PASS | 45.8s | $0.33 |
| 01_standalone_sdk/41_task_tool_set.py | ✅ PASS | 27.8s | $0.03 |
| 01_standalone_sdk/42_file_based_subagents.py | ✅ PASS | 55.2s | $0.05 |
| 01_standalone_sdk/44_model_switching_in_convo.py | ✅ PASS | 11.8s | $0.01 |
| 01_standalone_sdk/45_parallel_tool_execution.py | ✅ PASS | 3m 46s | $0.55 |
| 01_standalone_sdk/46_agent_settings.py | ✅ PASS | 11.6s | $0.01 |
| 01_standalone_sdk/47_defense_in_depth_security.py | ✅ PASS | 3.9s | $0.00 |
| 01_standalone_sdk/48_conversation_fork.py | ✅ PASS | 20.2s | $0.01 |
| 01_standalone_sdk/49_switch_llm_tool.py | ✅ PASS | 8.7s | $0.04 |
| 01_standalone_sdk/50_async_cancellation.py | ✅ PASS | 14.4s | $0.00 |
| 01_standalone_sdk/51_agent_hooks/main.py | ✅ PASS | 51.8s | $0.06 |
| 01_standalone_sdk/52_dynamic_workflow.py | ✅ PASS | 4m 1s | $0.17 |
| 01_standalone_sdk/53_client_defined_tools.py | ✅ PASS | 13.2s | $0.01 |
| 01_standalone_sdk/54_goal_completion_loop.py | ✅ PASS | 31.4s | $0.03 |
| 01_standalone_sdk/55_persistent_memory.py | ✅ PASS | 19.1s | $0.02 |
| 01_standalone_sdk/56_structured_output.py | ✅ PASS | 44.6s | $0.04 |
| 01_standalone_sdk/57_prompt_hooks/main.py | ✅ PASS | 12.1s | $0.00 |
| 02_remote_agent_server/01_convo_with_local_agent_server.py | ✅ PASS | 36.1s | $0.02 |
| 02_remote_agent_server/02_convo_with_docker_sandboxed_server.py | ✅ PASS | 1m 52s | $0.05 |
| 02_remote_agent_server/03_browser_use_with_docker_sandboxed_server.py | ✅ PASS | 1m 39s | $0.07 |
| 02_remote_agent_server/04_convo_with_api_sandboxed_server.py | ✅ PASS | 2m 4s | $0.03 |
| 02_remote_agent_server/06_custom_tool/main.py | ✅ PASS | 6m 41s | $0.05 |
| 02_remote_agent_server/07_convo_with_cloud_workspace.py | ✅ PASS | 48.7s | $0.03 |
| 02_remote_agent_server/08_convo_with_apptainer_sandboxed_server.py | ✅ PASS | 5m 29s | $0.03 |
| 02_remote_agent_server/09_acp_agent_with_remote_runtime.py | ✅ PASS | 1m 17s | $0.45 |
| 02_remote_agent_server/10_cloud_workspace_share_credentials.py | ✅ PASS | 35.4s | $0.04 |
| 02_remote_agent_server/11_conversation_fork.py | ✅ PASS | 38.1s | $0.00 |
| 02_remote_agent_server/12_settings_and_secrets_api.py | ✅ PASS | 2m 14s | $0.01 |
| 02_remote_agent_server/13_workspace_get_llm.py | ✅ PASS | 20.2s | $0.00 |
| 02_remote_agent_server/14_client_defined_tools.py | ✅ PASS | 27.0s | $0.02 |
| 02_remote_agent_server/15_openai_compatible_gateway.py | ✅ PASS | 40.8s | $0.01 |
| 02_remote_agent_server/16_deferred_init.py | ✅ PASS | 14.2s | $0.01 |
| 04_llm_specific_tools/01_gpt5_apply_patch_preset.py | ✅ PASS | 36.9s | $0.02 |
| 04_llm_specific_tools/02_gemini_file_tools.py | ✅ PASS | 56.2s | $0.10 |
| 05_skills_and_plugins/01_loading_agentskills/main.py | ✅ PASS | 16.1s | $0.02 |
| 05_skills_and_plugins/02_loading_plugins/main.py | ✅ PASS | 19.2s | $0.02 |
| 05_skills_and_plugins/04_mixed_marketplace_skills/main.py | ✅ PASS | 8.7s | $0.00 |
❌ Some tests failed
Total: 68 | Passed: 67 | Failed: 1 | Total Cost: $3.98
Failed examples:
- examples/01_standalone_sdk/28_ask_agent_example.py: Exit code 1
🧪 Integration Tests ResultsOverall Success Rate: 95.3% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_deepseek_deepseek_v4_flash
Skipped Tests:
litellm_proxy_anthropic_claude_sonnet_4_6
Failed Tests:
litellm_proxy_gemini_3.1_pro_preview
Failed Tests:
litellm_proxy_openai_gpt_5.5
litellm_proxy_minimax_MiniMax_M2.7
Skipped Tests:
|
🧪 Integration Tests ResultsOverall Success Rate: 100.0% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_deepseek_deepseek_v4_flash
litellm_proxy_anthropic_claude_sonnet_4_6
litellm_proxy_gemini_3.1_pro_preview
litellm_proxy_openai_gpt_5.5
litellm_proxy_minimax_MiniMax_M2.7
|
🧪 Integration Tests ResultsOverall Success Rate: 96.0% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_deepseek_deepseek_v4_flash
litellm_proxy_anthropic_claude_sonnet_4_6
litellm_proxy_gemini_3.1_pro_preview
Failed Tests:
Violations of Evaluation Criteria:The evaluation criteria explicitly states: "The agent must... Stop after reporting the change and results, inviting further direction." The agent violated this by continuing with substantial over-verification AFTER the first successful test run:
Impact Assessment:
Conclusion:While the technical outcome is correct and tests pass, the agent failed to follow the explicit instruction to stop after reporting results and inviting further direction. This represents a behavioral pattern of over-verification that doesn't align with the stated expectations for efficiency. (confidence=0.72) (Cost: $0.25) litellm_proxy_openai_gpt_5.5
litellm_proxy_minimax_MiniMax_M2.7
|
Release v1.44.0
This PR prepares the release for version 1.44.0.
Started by: @all-hands-bot
Release Checklist
integration-test)behavior-test)test-examples)security-scan)release-note-requiredPRs are accurately called out in the final release notesWhat happens on merge
When this PR is merged, the
create-release.ymlworkflow will automatically:v1.44.0and auto-generated notes, plus an explicit preamble for mergedrelease-note-requiredPRspypi-release.ymlto publish all packages to PyPIversion-bump-prs.ymlto create downstream version bump PRs🐳 Agent Server images for this PR — GHCR package, pull/run commands, and all pushed tags (click to expand)
• GHCR package: https://github.com/OpenHands/agent-sdk/pkgs/container/agent-server
Variants & Base Images
eclipse-temurin:17-jdknikolaik/python-nodejs:python3.13-nodejs22-slimgolang:1.21-bookwormPull (multi-arch manifest)
# Each variant is a multi-arch manifest supporting both amd64 and arm64 docker pull ghcr.io/openhands/agent-server:07a7fa1-pythonRun
All tags pushed for this build
About Multi-Architecture Support
07a7fa1-python) is a multi-arch manifest supporting both amd64 and arm6407a7fa1-python-amd64) are also available if needed