About Junji
We are committed to providing cost-effective artificial intelligence services.
Requests and responses sent through the Junji public API are used only to complete the current inference. We do not retain their contents or use them to train, fine-tune, or improve models. We keep only the minimum information required for billing, security, and reliable operation.
Public model API
We provide a public model API for developers and teams. Customers can use dependable model capability at a clear price without buying machines or maintaining an entire inference stack. We describe the available models, context capacity, and billing rules plainly, and our listed rates will not exceed current official prices.
Dedicated model access
For teams that need stable, long-term access to a particular model, Junji provides a dedicated service tailored to their needs. We deploy and operate an independent inference service and expose the agreed model, version, and service capacity through a stable API, without requiring the customer to maintain the full stack.
Private deployment and customization
For organizations with data-security, internal-process, or existing-hardware requirements, we provide private deployment and customization. The model runs in the customer’s own environment, where its data and systems remain under the customer’s control. We also tune the inference system for the actual hardware and workload to improve throughput, latency, reliability, and operating cost.
Cross-hardware and software-stack integration
Junji can integrate inference systems across hardware platforms and software stacks, including NVIDIA CUDA, Ascend, Hygon, AMD, and Intel. Our full-stack services extend from operators to driver software and cover inference runtimes, PyTorch compatibility layers, and framework integration, enabling end-to-end integration and optimization for each customer’s existing hardware, models, and business systems.