Black Sheep AI

SERVICES

Frontier AI, deployed where you decide

We publish the research, build the Bounded AI Architecture (Paddock, Ram and Watchman), and deploy production systems into infrastructure you control. No reselling someone else’s tools, no handoffs. Take one engagement or the whole pipeline.

Engagements

Six ways to work with us

Each engagement is a concrete deliverable, not a retainer. Every Ram model is benchmarked the way our published RAM evaluation was, across 7 model families and more than 40,000 questions, and Watchman's provenance detection is validated separately across five model families.

  • 01

    Ram Reasoning Models

    We compress a frontier model by 50–60% with no measurable quality loss, reasoning kept and certified. Tell us the device and the memory ceiling; we hand back a Ram model built to that budget, ready to bind to your Knowledge Plane. A 62GB model returns at 31GB, small enough for one GPU instance or a Mac Studio. Ten hours a day, that instance runs under $9k/year on demand and under $3k on spot, against $50–100k+ in API fees. Your weights stay in your infrastructure.

  • 02

    Watchman Governance

    Watchman is the governance layer of the Knowledge Plane: every module permission logged, every bound model verified. Before a third-party or compressed model enters your registry, we prove it is what it claims to be. Watchman detects and classifies weight modifications in minutes, works on the compressed releases people distribute, and returns evidence-grade reports with CycloneDX AI-BOM attestations mapped to the regulations arriving now. Every model gets audited before it goes live.

  • 03

    Paddock Knowledge Plane

    Your manuals, policies, and records, held in permissioned Knowledge Modules on infrastructure you control. Paddock returns exact citations and table-precise answers, with the reasoning model of your choice bound on top. It runs behind your firewall, and nothing leaves your boundary.

  • 04

    Knowledge Transfer & Domain Adaptation

    We fit a model to your domain, teaching it your terminology, your documents, and your tasks, then hand back a model that runs where you run it. When your knowledge belongs in the weights rather than a retrieval layer, we put it there. You own the result outright.

  • 05

    Deployment via Shepherd

    Shepherd is the platform that ties the pipeline together. It takes a frontier model, compresses it into a Ram reasoning model, certifies it with Watchman, binds it to your Paddock, and deploys it to your VPC, your Mac fleet, or an air-gapped facility. CI/CD integration is included, and nothing ships unless it clears the quality gate.

  • 06

    Managed Operations

    We keep the system running after it ships: monitoring and incident response, model updates with a fresh Watchman audit on each one, security patches, and compliance that holds up. 24/7 support from the people who built the stack.

End to end

Compress, deploy, operate

Most consultancies stop at advice. We take a frontier model end to end, from the research bench to a running system inside your perimeter.

  1. A press compressing a metal block
    01

    Compress

    Tell us your device and memory target. We fit the model to that budget and deliver it with a per-tensor quality report, so you know what the compressed model kept instead of guessing.

  2. A server being installed in a rack
    02

    Deploy

    Production deployment into your AWS, Azure, or GCP account, onto a Mac Studio or M-series fleet, into your data centre, or inside an air-gapped facility. Every token is processed within your security perimeter.

  3. An operations room with monitors
    03

    Operate

    Monitoring, model updates, and Watchman re-certification on every change, plus continuous improvement. We keep your private AI running at full performance long after handover.

DELIVERABLES

What each engagement delivers

Concrete outputs, and where they run. If a claim needs proof, the proof ships with it.
EngagementWhat you getRuns on
Ram Reasoning ModelsA Ram model built to your memory target, 50–60% smaller, reasoning kept, with a per-tensor quality reportAWS, Azure, GCP, or Apple Silicon
Watchman AuditEvidence-grade report and CycloneDX AI-BOM attestation, with modifications classified and locatedOn-prem, air-gapped, or in CI
PaddockThe Knowledge Plane: answers with exact citations and table-precise lookups, knowledge held in permissioned modulesYour hardware, behind your firewall
Knowledge TransferA model adapted to your domain, delivered as weights you ownWherever you deploy
ShepherdAutomated compress, audit, and deploy pipeline with quality gates and CI/CD integrationYour VPC, Mac fleet, or air-gapped facility
Managed Operations24/7 monitoring, model updates with re-certification, security patches, and complianceYour environment

Talk to our team

Bring us the model, the hardware, and the constraint. We will tell you what fits and what it costs. A 30-minute call, no sales pitch, just engineers.