Train locally. Prove it on Apple Silicon. Scale when needed.
A unified, native MLX stack for language, vision-language, embeddings, audio-language, and speech models on Apple Silicon.
Important
mlx-one is under active development. Native architecture support and model qualification are tracked separately. A family marked Architecture ✅ has a validated config, MLX model structure, registry entry, strict weight contract, and synthetic execution tests. It does not automatically mean every checkpoint, task, precision, or device has been qualified.
MLX is an excellent foundation for machine learning on Apple Silicon, but model workflows are often split across separate packages and incompatible task APIs. mlx-one is building those pieces as one coherent stack:
Python API + CLI
↓
Tasks: generation · embeddings · VLM · ASR · training
↓
Native model families + shared processors + safe loading
↓
Evaluation · benchmarking · evidence · qualification
↓
MLX on Apple Silicon
Click the architecture diagram to view it at full resolution.
The project owns its supported model math directly. Hugging Face is used as the
artifact ecosystem for configuration, tokenizer/processor assets, and safe
safetensors checkpoints—not as a remote-code execution runtime.
- Native MLX architectures across language, vision-language, embeddings, audio-language, and ASR.
- Shared attention, GQA, RoPE, normalization, feed-forward, MoE, vision, hybrid convolution, masking, and KV-cache primitives.
- Lazy metadata-driven model registration without initializing Metal during lightweight package imports.
- Strict configuration validation and deterministic weight contracts that reject unknown or shape-incompatible tensors.
- Safe local and pinned Hugging Face artifact inspection.
- Native Whisper loading, audio preprocessing, tokenization, decoding, language detection, segment timestamps, word timestamps, and WER/CER evaluation.
- Unified native text loading, tokenizer/chat templates, MLX 4-bit checkpoints, cached greedy and seeded sampling, stop sequences, token streaming, and serving for GPT-2, Qwen2/Qwen3, LFM2, OpenELM, and text-only Qwen3.5 families.
- Native Qwen3 embedding and reranking with byte-level BPE, safe loading, padding-aware batching, Matryoshka dimensions, normalization, cosine similarity, yes/no pair scoring, and stable ranking.
- Dataset validation, memory planning, LoRA/QLoRA SFT workflows, adapter export, evaluation, comparison, benchmarking, and evidence records.
- Backend-free tests plus opt-in Metal and real-checkpoint integration gates.
mlx-one requires Python 3.10 or newer and an Apple Silicon Mac for model execution.
git clone https://github.com/SSusantAchary/mlx-one.git
cd mlx-one
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .Install development tools with:
python -m pip install -e ".[dev]"FFmpeg is required only when the Whisper API receives an audio file. Already decoded mono 16-kHz waveforms can be passed directly from Python.
Load one supported text model and start the bundled local UI:
python -m pip install 'mlx-one[server]'
mlx-one serve mlx-community/Qwen3.5-0.8B-4bitClick the thumbnail to open the full-size UI screenshot.
To expose a shorter model name and protect /v1/* with a Bearer key:
mlx-one serve mlx-community/LFM2-350M-4bit \
--host 127.0.0.1 \
--port 8181 \
--alias lfm2-local \
--api-key local-secretEnter local-secret in the UI's API key field and select Connect. Terminal clients must
send the same key and use the alias as the request model:
curl http://127.0.0.1:8181/v1/models \
-H 'Authorization: Bearer local-secret'
curl -N http://127.0.0.1:8181/v1/chat/completions \
-H 'Authorization: Bearer local-secret' \
-H 'Content-Type: application/json' \
-d '{"model":"lfm2-local","messages":[{"role":"user","content":"Explain MLX in one sentence."}],"stream":true}'Authentication is disabled when neither --api-key nor MLX_ONE_API_KEY is configured. A
401 Unauthorized response from /v1/* means the Bearer key is missing or invalid; /health
and the static UI remain public.
The native server also supports model aliases, optional Bearer authentication, effective context
limits, request deadlines, cooperative parallel slots, prompt/context caching, context shifting,
reasoning output, quantized KV caches, and Qwen3.5 MTP speculative decoding. Run
mlx-one serve --help for the complete flag list.
Then open http://127.0.0.1:8080. The same process exposes an OpenAI-compatible API at
http://127.0.0.1:8080/v1, including streaming chat completions. Model execution remains inside
mlx-one's native model, tokenizer, sampling, generation, and MLX runtime. See
the Web UI and server guide for API and development details.
Check the machine and inspect a model without loading its weights:
mlx-one doctor
mlx-one inspect Qwen/Qwen2.5-1.5BBuild a workload-aware memory plan:
mlx-one plan inference \
--model Qwen/Qwen2.5-1.5B \
--hardware m4-air-32gb \
--precision bf16 \
--context-length 2048Generate text through the native GPT-2 stack:
mlx-one generate openai-community/gpt2 "The future of local AI is" \
--max-tokens 64 \
--temperature 0.8 \
--top-p 0.95 \
--seed 0Use --stream for incremental text, --json-output for a structured result,
or --offline with an existing local/cache copy. Model downloads occur only
when a user explicitly supplies an online repository without --offline.
Transcribe audio through the native Whisper stack:
mlx-one transcribe openai/whisper-tiny recording.m4a \
--language en \
--word-timestamps \
--json-outputOmit --language for automatic detection. Add --task translate for
speech-to-English translation or --offline to require local/cached assets.
Create document or instructed query embeddings:
mlx-one embed Qwen/Qwen3-Embedding-0.6B "A document to index" \
--dimensions 512 \
--json-output
mlx-one embed Qwen/Qwen3-Embedding-0.6B "local Apple Silicon training" \
--input-type queryRerank candidate documents while preserving their original indexes:
mlx-one rerank Qwen/Qwen3-Reranker-0.6B \
"Which Macs support MLX?" \
"MLX is designed for Apple silicon." \
"CUDA targets NVIDIA GPUs." \
--top-k 1 \
--json-outputFor reproducible evaluation and qualification, add --revision with an exact
model commit. Raw commit hashes are kept in integration tests and evidence
records rather than introductory examples.
| Registry type | Family | Native components | Architecture | Qualification |
|---|---|---|---|---|
gpt2 |
GPT-2 124M, 355M, 774M, 1.5B | Learned positions, fused QKV, byte BPE, safe loading, cache, generation and streaming | ✅ | Candidate; synthetic validation only |
llama |
Llama, MiniCPM5 1B/2B | GQA, explicit head dimensions, RoPE, 4-bit MLX loading, multi-EOS generation | ✅ | MiniCPM5 pinned integration gate |
qwen2 |
Qwen2, Qwen2.5, Qwen2.5-Coder | Dense Transformer, GQA, RoPE, cache | ✅ | Per checkpoint |
qwen3 |
Qwen3 dense | Bias-free attention, Q/K norm, explicit head dimensions | ✅ | Candidate |
qwen2_moe |
Qwen2-MoE | Top-k experts, shared expert, router outputs | ✅ | Architecture only |
openelm |
OpenELM 270M–3B | Layer-wise heads/FFN widths, fused QKV, GQA | ✅ | Verify release |
lfm2 |
LFM2 and LFM2.5 | Hybrid convolution/attention decoder | ✅ | Per checkpoint |
lfm2_moe |
LFM2 MoE | Hybrid decoder and sparse expert routing | ✅ | Excluded from ≤3B qualification |
| Registry type | Family | Native components | Architecture | Qualification |
|---|---|---|---|---|
qwen2_vl |
Qwen2-VL | Vision Transformer, 3D patches, merger, multimodal RoPE, image/video token insertion | ✅ | Verify exact checkpoint and processor |
qwen2_5_vl |
Qwen2.5-VL-3B | Window/full vision attention, biased SwiGLU vision tower, merger, multimodal RoPE | ✅ | Architecture only; verify checkpoint and processor |
qwen3_5 |
Qwen3.5 0.8B/2B | Hybrid gated-delta/full-attention text tower, learned/interpolated vision positions, multimodal RoPE | ✅ | Architecture only; synthetic validation |
lfm2_vl |
LFM2.5-VL | Vision tower, pixel unshuffle/projector, hybrid language model | ✅ | Verify exact checkpoint and processor |
| Registry type | Family | Native components | Architecture | Qualification |
|---|---|---|---|---|
bert |
all-MiniLM-L6-v2 | BERT encoder, mean pooling, normalization, cosine | ✅ | Candidate |
mpnet |
all-mpnet-base-v2 | MPNet relative positions, mean pooling, normalization, cosine | ✅ | Candidate |
qwen3_embedding |
Qwen3-Embedding-0.6B | Final-token pooling, instructed queries, Matryoshka dimensions, normalization, cosine | ✅ | Candidate; synthetic validation only |
qwen3_reranker |
Qwen3-Reranker-0.6B | Official pair prompt, yes/no logit scoring, probabilities, stable ranking | ✅ | Candidate; synthetic validation only |
lfm2_colbert |
LFM2/LFM2.5 ColBERT | Token embeddings, masks, late-interaction MaxSim | ✅ | Candidate |
| Registry type | Family | Native components | Architecture | Qualification |
|---|---|---|---|---|
whisper |
tiny, base, small, medium, large-v3, turbo | Encoder-decoder, safe loading, log-Mel, BPE, decoding, language, timestamps, WER/CER | ✅ | Candidate; pinned tiny/turbo smoke gates pass |
lfm2_audio |
LFM2.5-Audio | Audio encoder, Conformer, Depthformer, detokenizer, feature insertion | ✅ | Verify processor and codec weights |
The detailed family backlog and exact status are maintained in model_list.txt. Package ownership and dependency boundaries are defined in PROJECT_STRUCTURE.md.
mlx-one deliberately avoids turning one successful test into a broad support claim.
| Level | Meaning |
|---|---|
| Architecture ✅ | Config, model math, registry, weight contract, and synthetic tests pass |
| Candidate | The family or checkpoint still needs full integration evidence |
| Integration-tested | An exact checkpoint revision passes its task contract |
| Hardware-verified | A checksummed workload passes on a recorded Apple Silicon profile |
| Quality-verified | A pinned evaluation protocol passes its quality threshold |
Every evidence claim is scoped to a model revision, operation, precision, workload, software environment, and hardware profile.
mlx-one inspect MODEL [--revision REVISION] [--offline] [--json-output]
mlx-one hardware list
mlx-one hardware detect --json-output
mlx-one plan inference --model MODEL --hardware PROFILE
mlx-one plan train --model MODEL --hardware PROFILE --method autoInspection is metadata-only: it does not execute remote model code or open pickle checkpoints.
mlx-one generate MODEL PROMPT \
[--revision REVISION] [--offline] [--cache-dir DIRECTORY] \
[--max-tokens N] [--temperature T] [--top-k K] [--top-p P] \
[--seed N] [--stop TEXT] [--stream] [--json-output]GPT-2 uses the native path for direct generation, evaluation, and inference benchmarking. Native GPT-2 training and adapters are intentionally rejected until their own implementation and validation gates are complete.
mlx-one embed MODEL TEXT... \
[--input-type query|document] [--instruction TEXT] [--dimensions N] \
[--max-length N] [--batch-size N] [--offline] [--json-output]
mlx-one rerank MODEL QUERY DOCUMENT... \
[--instruction TEXT] [--top-k N] [--max-length N] [--batch-size N] \
[--offline] [--json-output]Both commands use the native Qwen3 tokenizer and strict safetensors loader. Document embeddings are unprefixed; query embeddings use the documented Qwen3 instruction format. Reranking returns the original document index, raw yes/no logit difference, and two-class probability.
mlx-one data validate \
--dataset examples/text-sft/train.jsonl \
--layout prompt-completionSupported layouts are instruction, messages, prompt-completion, and
text. A tokenizer can be supplied for truncation and loss-mask previews.
mlx-one train \
--model Qwen/Qwen2.5-Coder-1.5B-Instruct \
--revision PINNED_REVISION \
--dataset examples/text-sft/train.jsonl \
--config examples/text-sft/train.yaml
mlx-one export \
--checkpoint runs/qwen-coder-sft/adapters/mlx-one-checkpoint.json \
--output artifacts/qwen-coder-sftTraining is experimental. Start with the planner and pin the model revision.
mlx-one evaluate \
--model Qwen/Qwen2.5-Coder-1.5B-Instruct \
--revision PINNED_REVISION \
--dataset examples/text-sft/eval.jsonl \
--predictions examples/text-sft/predictions-smoke.json \
--runs-dir runs \
--json-output
mlx-one compare runs/base/result.json runs/candidate/result.json \
--gate examples/text-sft/quality-gate.yaml \
--format markdown
mlx-one benchmark inference \
--model Qwen/Qwen2.5-Coder-1.5B-Instruct \
--revision PINNED_REVISION \
--prompts examples/text-sft/prompts.jsonASR evaluation uses the pinned whisper-wer-v1 and whisper-cer-v1 profiles.
mlx-one registry list
mlx-one registry validate registry.json
mlx-one evidence publish --result runs/example/result.json --output evidence.jsonInspect and plan without initializing Metal:
from mlx_one import (
Precision,
WorkloadKind,
WorkloadSpec,
estimate_memory,
inspect_model,
load_hardware_profile,
)
model = inspect_model("Qwen/Qwen2.5-1.5B").model
hardware = load_hardware_profile("m4-air-32gb")
workload = WorkloadSpec(
kind=WorkloadKind.INFERENCE,
precision=Precision.BF16,
context_length=2048,
)
estimate = estimate_memory(model, hardware, workload)
print(estimate.to_json())Transcribe a file or mono 16-kHz waveform:
from mlx_one import WhisperDecodeOptions, transcribe
result = transcribe(
"openai/whisper-tiny",
"recording.wav",
language="en",
word_timestamps=True,
options=WhisperDecodeOptions(seed=0),
)
print(result.text)
for segment in result.segments:
print(segment.start, segment.end, segment.text)Embed and rerank through typed native retrieval results:
from mlx_one import embed, rerank
query = embed(
"Qwen/Qwen3-Embedding-0.6B",
"How does MLX use unified memory?",
input_type="query",
dimensions=512,
)
ranking = rerank(
"Qwen/Qwen3-Reranker-0.6B",
"How does MLX use unified memory?",
["MLX arrays share CPU and GPU memory.", "A recipe for sourdough."],
top_k=1,
)
print(query.embeddings[0])
print(ranking.items[0].index, ranking.items[0].score)Train through the typed SFT contract:
from mlx_one import SFTTrainer, TrainConfig
trainer = SFTTrainer(
model="Qwen/Qwen2.5-Coder-1.5B-Instruct",
revision="PINNED_REVISION",
train_dataset="examples/text-sft/train.jsonl",
args=TrainConfig(
output_dir="runs/qwen-coder-sft",
max_seq_length=512,
method="lora",
max_steps=10,
),
)
result = trainer.train()
print(result.to_json())- Configuration determines architecture; model-name conditionals do not.
- Shared operations are implemented once and reused across families.
- Family-specific classes stay inside their model packages.
- Weight mapping and sanitization are explicit and deterministic.
- Unsupported settings and unknown tensors fail visibly.
- Forward execution, generation, processing, and training remain separate layers.
- Training and inference converge on the same native model implementation.
- Lightweight inspection and registry imports do not initialize Metal.
- Qualification is based on reproducible evidence, never inference from a model name or parameter count.
The current repository gates include:
- Backend-free tests covering schemas, configs, registries, weight contracts, tokenization, processing, evaluation, loading policy, and CLI behavior.
- Opt-in Apple Silicon/Metal tests covering native model execution, shapes, caches, masks, multimodal feature insertion, embeddings, reranking, MoE routing, and ASR components.
- Pinned real-checkpoint smoke gates for
openai/whisper-tinyandopenai/whisper-large-v3-turbo. - Whisper log-Mel comparison within
1e-5against the pinned reference path. - Ruff, source/wheel builds, and package metadata checks.
These counts describe the current development tree and will change as coverage expands.
Run backend-free checks:
ruff check .
pytest -q
python -m build --no-isolation
python -m twine check dist/*Run native MLX execution tests on an Apple Silicon host:
MLX_ONE_RUN_MLX_TESTS=1 pytest -qRun pinned Whisper integration gates:
MLX_ONE_RUN_WHISPER_INTEGRATION=1 pytest -q \
tests/test_native_whisper_integration.pyOrdinary tests are network-free. Integration tests may download only the pinned assets needed by their explicit gate.
The implementation proceeds by evidence-backed vertical slices:
- Complete checkpoint loading, generation, and parity qualification for native language families.
- Expand the native retrieval APIs beyond the initial Qwen3 embedding and reranking vertical slice and qualify exact checkpoint revisions.
- Add complete image processing, generation, OCR/VQA evaluation, and training paths for native VLMs.
- Qualify Whisper revisions and expand native ASR, alignment, audio-language, and TTS coverage.
- Add native quantization, advanced training, export, and community evidence workflows.
Contributions should include the smallest relevant config, synthetic execution, weight-contract, integration, and documentation updates. Read CONTRIBUTING.md before opening a pull request. Report security issues through SECURITY.md, not a public issue.
Released under the MIT License.
