Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
quantumaikr
/
quant.cpp
Public
Notifications
You must be signed in to change notification settings
Fork
44
Star
399
Code
Issues
11
Pull requests
0
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Actions: quantumaikr/quant.cpp
Actions
All workflows
Workflows
CI
CI
Deploy WASM Demo to GitHub Pages
Deploy WASM Demo to GitHub Pages
Publish quantcpp to PyPI
Publish quantcpp to PyPI
Release
Release
Show more workflows...
Management
Caches
Deployments
All workflows
All workflows
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Showing runs from all workflows
will be ignored since log searching is not yet available
671 workflow runs
671 workflow runs
Workflow
Filter by Workflow
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching workflows.
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
feat: tq_batched_matmul_q4 — foundation for batched prefill
CI
#518:
Commit
ed4b087
pushed by
unamedkr
2m 37s
main
main
2m 37s
View workflow file
bench: prefill throughput script + 40× gap discovery
CI
#517:
Commit
84ea97c
pushed by
unamedkr
2m 41s
main
main
2m 41s
View workflow file
docs: throughput benchmark vs llama.cpp + session perf summary
CI
#516:
Commit
012ed22
pushed by
unamedkr
2m 44s
main
main
2m 44s
View workflow file
perf(q4k): 2-row ILP in q4k_int_dot_worker
CI
#515:
Commit
aadd059
pushed by
unamedkr
2m 30s
main
main
2m 30s
View workflow file
fix(metal): Phi-3.5 Q4_K_M garbage output under default build
CI
#514:
Commit
30dca7a
pushed by
unamedkr
4m 46s
main
main
4m 46s
View workflow file
test: regression guards for today's DeltaNet fix + long-seq stress + …
CI
#513:
Commit
361c40f
pushed by
unamedkr
4m 21s
main
main
4m 21s
View workflow file
feat(cli): filter chat-template control tokens from print_token
CI
#512:
Commit
8cea571
pushed by
unamedkr
10m 30s
main
main
10m 30s
View workflow file
feat(cli): --chat flag now correctly routes to per-family chat templates
CI
#511:
Commit
f140654
pushed by
unamedkr
5m 16s
main
main
5m 16s
View workflow file
test: relax Llama 3.1 8B check — raw '2+2=' is borderline
CI
#510:
Commit
1e8698b
pushed by
unamedkr
4m 34s
main
main
4m 34s
View workflow file
fix(server): revert Qwen3.5 think block injection (3/7 → 5/7)
CI
#509:
Commit
53386d2
pushed by
unamedkr
4m 55s
main
main
4m 55s
View workflow file
fix(qwen35): suppress <think> token — Qwen3.5-4B short prompts now wo…
CI
#508:
Commit
ba8a615
pushed by
unamedkr
4m 45s
main
main
4m 45s
View workflow file
fix(deltanet): restore L2 norm (removal causes output collapse) + dec…
CI
#507:
Commit
53b3323
pushed by
unamedkr
4m 39s
main
main
4m 39s
View workflow file
fix(deltanet): align decay formula with llama.cpp (#95)
CI
#506:
Commit
d26ca5e
pushed by
unamedkr
4m 34s
main
main
4m 34s
View workflow file
fix: restore Phi-3.5 as RLV default (Qwen3.5-4B: 6/20 on large-doc)
CI
#505:
Commit
7e2ca31
pushed by
unamedkr
5m 30s
main
main
5m 30s
View workflow file
feat: Qwen3.5-4B Acme 6/7 — auto-restart + model-agnostic prompts
CI
#504:
Commit
3ad0b80
pushed by
unamedkr
4m 40s
main
main
4m 40s
View workflow file
feat: model-agnostic RLV prompts — works with Phi-3.5, Qwen3.5, Qwen3
CI
#503:
Commit
13dc631
pushed by
unamedkr
4m 42s
main
main
4m 42s
View workflow file
analysis: Qwen3.5-4B scored 2/7 on Acme — Phi-3.5 remains best for RLV
CI
#502:
Commit
a64e8de
pushed by
unamedkr
5m 48s
main
main
5m 48s
View workflow file
fix(qwen35): early SSM probe fixes DeltaNet layer detection — Qwen3.5…
CI
#501:
Commit
e12fcbd
pushed by
unamedkr
4m 38s
main
main
4m 38s
View workflow file
feat(gemma4): E4B support + comprehensive numeric analysis
CI
#500:
Commit
b8a27d2
pushed by
unamedkr
5m 12s
main
main
5m 12s
View workflow file
feat: disable Qwen3 thinking mode by default (/no_think)
CI
#499:
Commit
bbb9159
pushed by
unamedkr
6m 6s
main
main
6m 6s
View workflow file
feat: quant_generate continues from loaded KV cache (#83)
CI
#498:
Commit
e273f2b
pushed by
unamedkr
5m 31s
main
main
5m 31s
View workflow file
fix(gemma4): numeric comparison with MLX-LM — divergence after layer 0
CI
#497:
Commit
2899fb8
pushed by
unamedkr
4m 43s
main
main
4m 43s
View workflow file
feat(default): switch to Qwen3-4B — 2.4x faster + best quality
CI
#496:
Commit
6ea3215
pushed by
unamedkr
4m 34s
main
main
4m 34s
View workflow file
fix(gemma4): token injection + confirmed forward pass bug
CI
#495:
Commit
91606ce
pushed by
unamedkr
4m 36s
main
main
4m 36s
View workflow file
fix: load_context resets chat state + KV cache saves 57 tokens (#83)
CI
#494:
Commit
ffe668f
pushed by
unamedkr
4m 39s
main
main
4m 39s
View workflow file
Previous
1
2
3
4
5
6
…
26
27
Next
You can’t perform that action at this time.