CoderAI is a multimodal, multi-backend local model orchestrator with a drop-in OpenAI-compatible API server, so you can run models on your own GPUs instead of someone else's cloud.
It supports NVIDIA (CUDA), AMD and Intel (Vulkan) backends with per-model configuration, on-demand load/unload, smart VRAMβRAMβdisk offloading, prompt caching and aggregation, parallel execution and a request queue.
A full Web Studio covers chat, image, video and audio generation β text-to-image/video, image-to-video, upscaling, inpainting, face-swap, TTS/STT, voice cloning, dubbing and custom multi-step pipelines.
Screenshots




