An AI-powered code reviewer that runs entirely on-device using Snapdragon NPUs - keeping proprietary source code private while delivering instant, offline code reviews.
Built for the Snapdragon® Multiverse Hackathon (Bangalore) - Top 50 Teams (India level), Phase 2 build.
After the build phase ended, our team was selected among the top 8 finalists out of all 50 teams and presented live on stage in the final battle - finishing 2nd overall, just 0.5 points behind 1st place.
- ✅ Runs entirely on Snapdragon NPUs - no cloud API, no CPU fallback needed
- ✅ Reviews Git commits and GitHub PRs automatically
- ✅ Never uploads a single line of source code
- ✅ Streams findings to a mobile app in real time for triage
- ✅ Lets the developer push back - the AI reconsiders any finding flagged as a false positive, given the developer's reasoning
- ✅ Generates a fully offline PDF report
git commit (or a GitHub PR) triggers diff extraction → a local LLM running
entirely on a Snapdragon NPU analyzes the diff → findings stream live to
a mobile triage app → the developer approves each finding or flags it as a
false positive → the AI reconsiders flagged findings given the developer's
reasoning → a final PDF report is generated on-device.
No internet dependency for review. No API keys. No code ever leaves the machine. Confirmed end-to-end on Snapdragon X Elite hardware, running fully on-NPU via Qualcomm AI Hub's GenieX runtime - not a CPU fallback, not a mock.
- Why DevMesh exists
- Team
- How it works
- Results
- Tech stack
- Repository structure
- Setup - from scratch
- Run & usage instructions
- Screenshots & demo
- Verifying the setup (tests)
- Known issues
- Notes
- Future work
- References
- License
AI code review tools (GitHub Copilot, CodeRabbit, Cursor, etc.) are great, but every one of them sends your source code to a cloud API. For a huge class of developers - fintech, healthtech, defense, government, or just companies with a strict data-residency policy - that's a non-starter. India alone has an estimated 5.8M developers, and a meaningful slice of them work somewhere that simply will not permit proprietary code to leave the building, let alone the country.
DevMesh answers a simple question: can you get a genuinely useful AI code reviewer without any of the code ever touching the network? The answer is yes, if the model runs on the same machine that owns the code - and modern Snapdragon NPUs are fast enough to make that practical rather than theoretical. DevMesh runs a 4B-parameter instruction-tuned model (Qwen3-4B-Instruct-2507) entirely on-NPU via Qualcomm AI Hub / GenieX, reviews your diff the moment you commit, and lets you triage the findings from your phone - with zero code ever transmitted anywhere.
| Name | Role | |
|---|---|---|
| Hardik Parmar | Team Lead - LLM integration, prompt engineering, response parsing, WebSocket server, backend orchestration, structured logging | hj.parmar1@tcs.com, hi@hardikjp7.com |
| Vatsal Bhavesh | React Native mobile app, WebSocket client, fpdf2 PDF report generator, false-positive UX | vb.mandaliya@tcs.com |
| Dhruv | Git hook, GitHub webhook listener, FastAPI server | ds.bailkur@tcs.com |
- Trigger. A
git commitfires thepost-commithook, or a GitHub PR webhook lands on the FastAPI listener. Both routes funnel into the same already-running backend process - never two independent pipeline instances (see Known Issues / architecture notes for why that matters). - Diff extraction.
git diff --function-contextgives each hunk full enclosing-function context, not just the default 3 lines. - Review. Each hunk is sent to the LLM with a fixed-format prompt.
On Snapdragon X Elite hardware this runs via
geniex inferon the NPU; Ollama is the CPU fallback for development machines without the Snapdragon backend. - Parsing. The model's raw text is parsed into structured
CRITICAL/MAJOR/MINOR/SUGGESTIONfindings. - Broadcast. Findings stream to the mobile app over a WebSocket, grouped by commit.
- Triage. On the phone, each finding is either Approved or marked False Positive (a real justification is required - a one-word dismissal is rejected before it ever reaches the model).
- Reiteration. When the developer asks to generate the report, every
false-positive-flagged finding is re-sent to the LLM along with the
original diff and the developer's comment. The model replies
MAINTAINED,WITHDRAWN, orPARTIALLY_VALIDwith a short explanation - so a developer's dismissal doesn't just silently make an issue vanish. - Report. A designed PDF (plus a plain-text debug copy) is written locally, covering every approved finding and every reiterated false positive.
Measured on Snapdragon X Elite CRD hardware, on-NPU, via GenieX:
| Metric | Value |
|---|---|
| Execution | 100% offline - zero network calls during review |
| Model | Qwen3-4B-Instruct-2507 (non-thinking variant) |
| Prefill throughput | ~1,301 tok/s |
| Decode throughput | ~23.1 tok/s |
| Observed per-hunk latency | ~11.8s – 34.4s (validated on Snapdragon X Elite hardware) |
| Context window | 4,096 tokens |
| Network dependency | None - code never leaves the device |
These numbers come from real per-hunk benchmark runs on venue hardware
(backend/benchmark.py), not simulated or CPU-only figures - see
Verifying the setup to reproduce them.
| Component | Technology |
|---|---|
| LLM model | Qwen3-4B-Instruct-2507 - selected after direct benchmarking on Snapdragon X Elite: native NPU/QAIRT runtime, 4096-token context, 1,301 tok/s prefill, 23.1 tok/s decode |
| LLM runtime (production) | geniex infer - Qualcomm AI Hub's GenieX CLI, one-shot per hunk, on-NPU |
| LLM runtime (dev/fallback) | Ollama (qwen3:4b-instruct), CPU-only |
| Backend | Python 3.10+, FastAPI, websockets |
| Git/PR trigger | Shell hook (hooks/post-commit) → FastAPI webhook listener |
| Mobile app | React Native (Expo) - display/triage only, zero on-device inference |
| Report generation | fpdf2 - pure-Python PDF rendering (no native deps, ARM64-safe) |
| Structured logging | Custom devlog.py - console + persistent commit-tagged log file |
devmesh/
├── backend/ Python - review pipeline, WebSocket server, webhook listener, PDF report generation
├── mobile/ React Native (Expo) - triage app
├── hooks/ post-commit git hook
├── samples/ demo buggy files used for live-review testing
├── setup.sh one-command install + hook install + start everything
├── README.md
└── LICENSE MIT
Full backend/ file layout (click to expand)
backend/
├── run_review.py CLI entrypoint - reviews one commit per invocation
├── webhook_server.py FastAPI listener - real GitHub PR + local-commit trigger
├── diff_extractor.py git diff, hunk splitting, commit metadata
├── prompt_builder.py fixed review-prompt template
├── llm_client.py review_hunk() - geniex / ollama / mock backend dispatch
├── response_parser.py tolerant finding parser (see Known Issues)
├── review_session.py in-memory session: finding ids, decisions, commit info
├── reiteration.py false-positive re-check prompt + parser
├── report_trigger.py orchestrates reiteration pass + report generation
├── report_generator.py writes PDF (via pdf_report.py) + plain-text debug copy
├── pdf_report.py fpdf2 PDF layout
├── ws_broadcaster.py bidirectional WebSocket server
├── devlog.py structured logging (console + logs/devmesh.log)
├── benchmark.py per-hunk latency benchmarking harness
└── requirements.txt
These steps take you from a clean machine to a running system. Two backend paths are supported - pick based on what hardware you have:
- Snapdragon X Elite (or similar ARM64 NPU device) → the real, on-NPU path via Qualcomm AI Hub / GenieX. This is what the hackathon submission runs on.
- Any other machine (Windows/Mac/Linux, x86 or otherwise) → the Ollama CPU fallback. Functionally identical pipeline, just slower and not running on an NPU.
| Tool | Why | Check |
|---|---|---|
| Python 3.10+ | Backend pipeline, webhook listener | python3 --version |
| Node.js (LTS) + npm | Mobile app (Expo/React Native) | node --version |
| Git | Version control | git --version |
| Expo Go (App Store / Play Store) | Run the mobile app on your phone | - |
| Either: Ollama (dev/fallback) | Local CPU LLM | ollama --version |
| Or: Qualcomm AI Hub / GenieX SDK (venue path) | On-NPU LLM inference | geniex --version |
git clone <this-repo-url> devmesh
cd devmeshcd backend
pip install -r requirements.txt
# On Linux with an externally-managed Python environment:
pip install -r requirements.txt --break-system-packages
# or use a virtualenv:
python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt
cd ..cd mobile
npm install
cd ..# Install from https://ollama.com, then:
ollama pull qwen3:4b-instruct
ollama serve # leave running in its own terminalDevMesh defaults DEVMESH_BACKEND to qnn (the NPU path). To use Ollama
instead, set the env var when running the backend:
export DEVMESH_BACKEND=ollama # Windows PowerShell: $env:DEVMESH_BACKEND="ollama"-
Install the Qualcomm AI Hub / QAIRT SDK and confirm
geniexis on yourPATH(geniex --version). -
Download or compile the Qwen3-4B-Instruct-2507 model for your device via Qualcomm AI Hub.
-
No extra env var is needed -
qnnis the default backend. Optional overrides:Variable Purpose Default DEVMESH_GENIE_MODELModel identifier passed to geniex inferai-hub-models/Qwen3-4B-Instruct-2507DEVMESH_GENIE_COMPUTECompute target npuDEVMESH_GENIE_THINKEnable/disable the model's thinking mode falseDEVMESH_GENIE_PROMPT_FILEPath to the fixed prompt file geniexreads frombackend/devmesh_geniex_prompt.txt
From the repository root:
chmod +x setup.sh
./setup.shsetup.sh will:
- Install backend Python dependencies (non-fatal on failure - installs the hook regardless).
- Install
hooks/post-commitinto.git/hooks/post-commit. - Start the FastAPI webhook listener in the background on port
8000(this also starts the WebSocket server on port8765). - Start the Expo mobile development server in the background on a fixed port (
5678), logging its QR code tologs/devmesh_expo.log.
Safe to re-run anytime—it won't double-start the webhook listener.
Note: If the Expo QR code does not appear after running
setup.sh, start the Expo development server manually:cd mobile npx expo startOnce the Expo CLI starts, the QR code should be displayed. If your physical device does not automatically appear as connected, use the Down Arrow (↓) key in the Expo CLI to select your device and approve the connection. After approval, you can scan the QR code (if needed) and continue development normally.
If you are not inside a Git repository yet, run
git initfirst. The script will display a warning and skip the hook installation otherwise.
Edit mobile/App.js and set SERVER_IP to your laptop's LAN IP (phone and
laptop must be on the same Wi-Fi):
const SERVER_IP = "x.x.x.x"; // your laptop's LAN IP, not the phone'sFind your IP with ipconfig (Windows) or ifconfig / ip addr
(Mac/Linux).
cat logs/devmesh_expo.logScan it with Expo Go on your phone. The app shows a live "Waiting for AI PC" status with a retry counter until a real review runs - there's also a Show Demo Data button for offline UI demoing, clearly banner-labeled as sample data.
A. Automatic, via a real commit:
git add samples/buggy_auth.py
git commit -m "test commit"The hook fires in the background; git commit returns immediately.
B. Simulate a GitHub PR:
curl -X POST http://localhost:8000/webhook \
-H "X-GitHub-Event: pull_request" \
-H "Content-Type: application/json" \
-d '{"action": "opened", "pull_request": {"number": 1}}'C. Manual CLI, for debugging one component at a time:
cd backend
python run_review.py # reviews the last commit
python run_review.py --staged # reviews staged changes instead
python run_review.py --repo /path/to/other/repoAll three converge on the same output: findings pushed to any connected
mobile client, and (once the developer resolves every finding on mobile and
taps Generate Report) a PDF written to backend/devmesh_report_<short_commit_hash>.pdf.
- Each finding shows severity, file/line, description, and a suggested fix.
- Tap Approve or Mark False Positive (the latter requires typing a real reason).
- Tap Generate Report once every finding is resolved. Any false positives without a verdict yet trigger an AI re-check first ("Verifying false positives…"), then the report generates automatically.
- A native alert confirms the report's file path when it's ready, and resets the screen for the next commit.
| Variable | Purpose | Default |
|---|---|---|
DEVMESH_BACKEND |
qnn (on-NPU, Snapdragon hardware) or ollama (CPU dev/fallback) |
qnn |
DEVMESH_MODEL |
Ollama model tag (only used when DEVMESH_BACKEND=ollama) |
qwen3:4b-instruct |
DEVMESH_GENIE_MODEL |
GenieX model identifier | ai-hub-models/Qwen3-4B-Instruct-2507 |
DEVMESH_GENIE_COMPUTE |
GenieX compute target | npu |
DEVMESH_GENIE_THINK |
GenieX thinking-mode toggle | false |
DEVMESH_MOCK_LLM |
1 to bypass the LLM entirely with a canned response (fast pipeline testing) |
0 |
DEVMESH_TIMEOUT |
LLM call timeout, seconds | 420 |
DEVMESH_LOG_LEVEL |
DEBUG / INFO / WARNING / ERROR |
INFO |
DEVMESH_LOG_DIR |
Where devmesh.log is written |
./logs |
GITHUB_TOKEN |
Optional, avoids GitHub's 60 req/hr unauthenticated rate limit for real PR diff fetches | unset |
The developer reviews findings, approves valid issues, or marks false positives and then click on generate report to generate PDF report.
A commit automatically triggers the on-device review pipeline.
Offline report summarizing approved findings and AI re-evaluation results.
Watch the complete workflow here:
These checks confirm each layer works independently, cheapest first:
cd backend
# 1. Parser sanity check (no LLM needed)
python response_parser.py
# Expect: prints several Finding(...) objects
# 2. Prompt template sanity check (no LLM needed)
python prompt_builder.py
# Expect: prints the filled-in prompt + measured token overhead
# 3. Full pipeline without any LLM (fastest end-to-end check)
DEVMESH_MOCK_LLM=1 python run_review.py
# Expect: instant run, canned findings broadcast + report written
# 4. Real LLM connection
python llm_client.py
# Expect: a latency number and raw model output resembling
# "[CRITICAL] file.py:1 - ... Fix: ..."
# 5. Webhook listener health
curl http://localhost:8000/health
# Expect: {"status":"ok","service":"devmesh-webhook-listener"}Latency numbers can be logged systematically with:
python benchmark.py # benchmarks the last commit
python benchmark.py --runs 3 # repeats each hunk 3x and averagesWrites timestamped CSV + JSON to backend/benchmark_results/.
| Symptom | Likely cause |
|---|---|
Could not reach Ollama at http://localhost:11434 |
ollama serve isn't running |
geniex executable not found |
GenieX/QAIRT SDK not installed or not on PATH |
git diff failed |
Not inside a git repo, or no commit exists yet |
Model output has no [SEVERITY] lines |
Model ignored the format - consider iterating the prompt wording in prompt_builder.py |
| Phone stuck on "Waiting for AI PC" | Check SERVER_IP in mobile/App.js, confirm same Wi-Fi, check for AP/client isolation |
| Nothing reaches the phone but the report file updates | Two processes are both trying to own the WebSocket port (8765) - only ever run one run_review.py/webhook-listener process at a time |
- No message backlog/replay - a mobile client that connects after a review already ran will miss those findings. Connect the mobile client before triggering a review.
- Real GitHub PR diffs use GitHub's default 3-line context (not the full-function context local commits get) - GitHub's diff API has no equivalent option.
- Single review session at a time, in-memory only - a process restart loses in-flight decisions. Sufficient for a single demo/review run; would need persistence for longer-lived use.
- Why every commit routes through one long-running process: two
independent backend processes would each hold their own isolated
session state, silently causing the mobile app to see stale/null commit
data depending on which process it happened to be connected to. DevMesh
enforces a pre-flight port check on its WebSocket port (
8765) that fails loudly rather than allowing a second instance to start. - Why the mobile app does no on-device inference: the privacy story is "the machine that owns the code also runs the model" - the phone is strictly a triage/display client over a local WebSocket, by design, not a limitation.
- Why fpdf2 instead of WeasyPrint/HTML+CSS: WeasyPrint pulls in native libraries (Cairo/Pango/GDK-PixBuf) that are a real risk on Windows ARM64. fpdf2 is pure Python - no native dependencies - which matters for "runs everywhere including the Snapdragon ARM64 laptop" as a hard requirement, not a preference.
- Structured logs live in
logs/devmesh.log(Python side, commit-tagged, stage-timed - a hang shows up as aSTAGE STARTline with no matchingSTAGE END) andlogs/devmesh_hooks.log(bash side).
- Multi-developer review sessions (persisted, not single in-memory session)
- Persistent review history across restarts
- VS Code / JetBrains extension for in-editor triage
- Incremental review caching - skip re-reviewing unchanged hunks
- Support for larger NPU-optimized models as Snapdragon NPU headroom grows
- Message backlog/replay for mobile clients that connect mid-review
- Qualcomm AI Hub - model benchmarking and compilation for Snapdragon NPUs
- Qualcomm AI Hub documentation - model compilation and deployment reference
- GenieX / QAIRT documentation -
Qualcomm AI Engine Direct (QAIRT) SDK, the runtime backing
geniex infer - Qwen3 model family - the LLM DevMesh runs on-device
- Qwen3 technical report - model architecture and training details
- Ollama - local LLM runtime used for CPU dev/testing
- fpdf2 - pure-Python PDF generation
- Expo / React Native - mobile app framework
Apache-2.0 license - see LICENSE.



