You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds DeputyDev v1 to the 50-PR offline code-review benchmark, including GPT-5.5 judged results.
What is included
one deputydev-v1 review entry for each of the 50 benchmark tasks
267 raw review comments
224 extracted candidates
GPT-5.5 candidates, dedup groups, and evaluations
DeputyDev registration in the evaluated-tools list and dashboard
support for collecting reviews posted through a human GitHub account
a focused downloader regression test
GPT-5.5 result
TP: 74
FP: 145
FN: 63
micro precision: 33.79%
micro recall: 54.01%
skipped evaluations: 0
judge errors: 0
Validation
exactly 50 DeputyDev v1 benchmark reviews
exactly 50 candidate and evaluation entries
43 dedup-group entries
only deputydev-v1 is present as the tool key in the GPT-5.5 result files
downloader tests: 8 passed
focused Ruff checks passed
generated dashboard includes the gpt-5.5 model and DeputyDev v1
The source clone repository names and PR URLs retain their original internal run identity for provenance, while the submitted public benchmark tool identity is deputydev-v1.
hey @ankit27755, can you provide the Github org of the repos/PRs used for your run of the benchmark?
Hey @ashleyzhang01 :
GitHub organization: https://github.com/dd-benchmarking
The organization also contains repositories from earlier internal benchmark runs. The 50 repositories used for the submitted DeputyDev V5 result are specifically those matching:
Repositories containing identifiers such as deputydev-v2, deputydev-v2-gpt-5-6, or other run dates were from previous experiments and were not used for the submitted V5 score. I can also provide the exact list of 50 PR URLs if needed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds DeputyDev v1 to the 50-PR offline code-review benchmark, including GPT-5.5 judged results.
What is included
deputydev-v1review entry for each of the 50 benchmark tasksGPT-5.5 result
Validation
deputydev-v1is present as the tool key in the GPT-5.5 result filesgpt-5.5model and DeputyDev v1The source clone repository names and PR URLs retain their original internal run identity for provenance, while the submitted public benchmark tool identity is
deputydev-v1.