Practical, plain-English guides to model compression, MLX on Apple Silicon, and AI governance, each linking to the research and products behind it.
mlx_lm.convert, bit widths, group size, and the pitfalls the docs skip.
The two FP4 formats, MLX support status, and why sub-4-bit quality is hard.
Router sensitivity, per-expert bits, and the MLX constraint that shapes your options.
Top-k gating, softmax scores, and why confidence scales inversely with expert count.
Why experts that look dead aren't, and what to measure instead.
Streaming shards, Metal buffer limits, and an honest note on the search term.
What triggers the ValueError, and four concrete fixes.
Predict RAM from bits-per-weight, and why decode is memory-bandwidth-bound.
What MMLU-Pro measures, and why a 31 GB compressed build matched full precision.
RAI as the controls that decide who ships and who can stop an AI system.
The four tiers, how a system gets classified, and why a flat checklist fails.
The specific triggers that make an assessment mandatory, and why.
Policy-layer content governance vs runtime LLM guardrails, and how they fit.
A RACI mapped to model owners, data stewards, and the three lines of defense.
Likelihood x impact x detectability scoring that feeds risk tiering.
Competency, authority, and the ethics-liaison role, beyond generic upskilling.