ChatterUI is a native mobile frontend for LLMs.
Run LLMs on device or connect to various commercial or open source APIs. ChatterUI aims to provide a mobile-friendly interface with fine-grained control over chat structuring.
If you like the app, feel free support me here:
Use on-device Models or APIs
Modify And Customize
Personalize Yourself
- Run LLMs on-device in Local Mode
- Connect to various APIs in Remote Mode
- Chat with characters. (Supports the Character Card v2 specification.)
- Create and manage multiple chats per character.
- Customize Sampler fields and Instruct formatting
- Integrates with your device’s text-to-speech (TTS) engine
This fork (akumaburn/ChatterUI) builds on upstream with the following changes:
- Runs a current build of llama.cpp via a vendored copy of the cui-llama.rn adapter (under
vendor/cui-llama.rn), compiled into the app instead of pulled from npm. - Exposes all model-load settings (Flash Attention, KV-cache K/V quantization, micro-batch, mmap/mlock, unified KV cache, full SWA cache, CPU MoE layers, and more) and the full set of llama.cpp samplers (including Top-n-sigma).
- Removes arbitrary slider min/max caps: numeric fields accept any valid value while the sliders keep a usable range.
- Ships roleplay-oriented sampler defaults (Min-P, a light repetition penalty, and DRY) with unlimited generation length by default (generation ends only on a stop token / EOS, context-shifting to continue).
- Fixes a New-Architecture (bridgeless) crash on React Native 0.83+ where the native llama module failed to install.
Download and install latest APK from this fork's releases page.
iOS is Currently unavailable due to lacking iOS hardware for development
ChatterUI uses llama.cpp under the hood to run gguf files on device. A custom adapter is used to integrate with react-native: cui-llama.rn. In this fork the adapter is vendored under vendor/cui-llama.rn (consumed as a local file: dependency) and tracks a current build of llama.cpp; npm install links it automatically and no extra fetch step is required.
To use on-device inferencing, first enable Local Mode, then go to Models > Import Model / Use External Model and choose a gguf model that can fit on your device's memory. The importing functions are as follows:
- Import Model: Copies the model file into ChatterUI, potentially speeding up startup time.
- Use External Model: Uses a model from your device storage directly, removing the need to copy large files into ChatterUI but with a slight delay in load times.
After that, you can load the model and begin chatting!
Note: For devices with Snapdragon 8 Gen 1 and above or Exynos 2200+, it is recommended to use the Q4_0 quantization for optimized performance.
Remote Mode allows you to connect to a few common APIs from both commercial and open source projects.
- koboldcpp
- text-generation-webui
- Ollama
- OpenAI
- Claude (with ability to use a proxy)
- Cohere
- Open Router
- Mancer
- AI Horde
- Generic Text Completions
- Generic Chat Completions
These should be compliant with any Text Completion/Chat Completion backends such as Groq or Infermatic.
Is your API provider missing? ChatterUI allows you to define APIs using its template system.
Read more about it here!
To run a development build, follow these simple steps:
- Install any Java 17/21 SDK of your choosing
- Install
android-sdkviaAndroid Studio - Clone the repo:
git clone https://github.com/Vali-98/ChatterUI.git
- Install dependencies via npm and run via Expo:
npm install
npx expo run:android
Requires Node.js, Java 17/21 SDK and Android SDK. Expo uses EAS to build apps which requires a Linux environment.
- Clone the repo.
- Rename the
eas.json.exampletoeas.json. - Modify
"ANDROID_SDK_ROOT"to the directory of your Android SDK - Run the following:
npm install
eas build --platform android --local
Currently in development