CoderAI

multimodal local model orchestrator

🌐 live

CoderAI is a multimodal, multi-backend local model orchestrator with a drop-in OpenAI-compatible API server, so you can run models on your own GPUs instead of someone else's cloud.

It supports NVIDIA (CUDA), AMD and Intel (Vulkan) backends with per-model configuration, on-demand load/unload, smart VRAM→RAM→disk offloading, prompt caching and aggregation, parallel execution and a request queue.

A full Web Studio covers chat, image, video and audio generation β€” text-to-image/video, image-to-video, upscaling, inpainting, face-swap, TTS/STT, voice cloning, dubbing and custom multi-step pipelines.

Screenshots

Live overview β€” NVIDIA and AMD backends side by side, with per-engine health
Live overview β€” NVIDIA and AMD backends side by side, with per-engine health
Per-model configuration β€” quantization, backends and pipeline components
Per-model configuration β€” quantization, backends and pipeline components
Model discovery and download straight from HuggingFace
Model discovery and download straight from HuggingFace
Live task view β€” generations and LoRA training, with GPU/VRAM telemetry
Live task view β€” generations and LoRA training, with GPU/VRAM telemetry
The local model library β€” storage, capabilities and per-model config
The local model library β€” storage, capabilities and per-model config

Tags

PythonCUDA / VulkanOpenAI-compatiblemultimodalself-hosted

Links

Back to all projects

$ ./work-with-me

Interested in this project, or need something like it built? I'm available for freelance & contract work.