AI product studio

Build AI-native products. Ship them with control.

NCM Labs helps teams turn LLMs and agents into products that actually work: prototypes in days, then hardened with evaluation gates, supervision and full traceability.

  • Our own language FORGE makes LLM uncertainty a type, not a surprise.
  • Agents in real tools From chat to Fusion 360 geometry and CNC G-code.
  • We run what we build Our in-house products run on the same methods.

Live demo · Agent console

Watch an agent team work. Then take the human seat.

Pick a job, inject a problem, press run. Every step shows the tool it called, how sure it was and which rule fired. When the stakes rise, the agents stop and ask you.

ncm://console/document-inbox Ready
1 · Choose a job

Request

“Close September for Example Studio SRL: collect every receipt and invoice, match the payments and hand it all to our accountant.”

2 · Inject a problem

Scripted replay with synthetic data. Same method we use on real systems.

Run trace

0 steps
  1. Press Run agents to start the trace.

Live artifact

Handoff pack
Documents for September
DocumentAmount Status
INV-0912 · Hosting49.00 EUR
RCPT-0915 · Supplies212.40 RON
INV-0918 · Contractor3,200.00 RON
RCPT-0921 · Fuel310.75 RON
INV-0926 · Software19.99 EUR
Bank · 24 Sep transfer1,450.00 RON
  • Collected
  • Extracted
  • Matched
  • Exceptions
  • Handoff sent
Elapsed
0.0 s
Cost
€0.000
Confidence
—
Gates passed
0

Services · What we build with you

The new tools, put to work on your product.

LLMs and agents change how fast software gets built and what it can do. We bring both, plus the engineering that keeps them honest.

AI product builds

From idea to a product people use: LLM features and agent-powered apps, designed around the moments where the model can be wrong.

A working product, not a demo.

LLM apps agents full stack
Discuss this

Agent workflow automation

Repetitive, document-heavy or operational work handed to supervised agents with budgets, confidence thresholds and a human sign-off where it matters.

Hours back every week, with an audit trail.

supervision human-in-the-loop tracing
Discuss this

LLM integration & evaluation

Choosing, adapting and serving the right model for your task, then proving it with held-out evaluations against explicit baselines.

Decisions backed by measurements.

model choice fine-tuning evals
Discuss this

AI for engineering & making

Connecting assistants to CAD, fabrication and embedded systems through inspectable tools and validation between intent and geometry.

Conversation becomes checked, physical output.

CAD CAM embedded
Discuss this

Process · Idea to production

Fast where it's cheap. Careful where it counts.

Agents compress the build. Gates, evaluations and human checkpoints are what make the result safe to rely on.

  1. 01

    Frame

    We map the workflow, where a model adds value, and where it must never act alone. You get a scoped plan and success metrics.

  2. 02

    Prototype in days

    Agents and LLMs let us put a working slice in your hands fast, on your real inputs, so decisions come from use rather than slides.

  3. 03

    Harden with gates

    Evaluation sets, confidence thresholds, budgets and escalation paths turn the prototype into something you can trust under load.

  4. 04

    Ship & operate

    Deployed, traced and measured. Every decision, cost and fallback leaves evidence your team can inspect and improve.

Work · Built in-house and in the open

We build the tools we sell with.

A programming language for LLM systems, bridges from agents into CAD and CNC, and products we run ourselves. Open-source where we can, measured everywhere.

Open source Programming language · Rust · Open source

FORGE

A programming language for oracle-augmented computation. LLM calls are oracle queries with their own type: uncertainty is tracked at compile time, deterministic code is structurally separated from stochastic code, and failure handling is declared, not bolted on.

Uncertain results become explicit, handled values.

classify_intent.forge
task classify_intent
  needs message: Text
  gives Intent

  do
    result = classify message into ["buy", "support", "cancel", "other"]

    when result.sure(above: 0.85) -> give result    when result.sure              -> give result with flag("low-confidence")    when result.unsure            -> give ask_for_clarification(message)    else                          -> escalate to human
Task

give

“support” returned to the caller.

In-house In-house product · In development

Document inbox for accountants

A portal that collects monthly accounting evidence, resolves missing documents and payment exceptions, and prepares a traceable handoff for the accountant. Extraction is measured against an evaluation corpus before it is trusted.

Month-end becomes a checklist, not a chase.

In-house In-house tooling · In development

LLM-native infrastructure control plane

One structured command surface an agent can drive to manage compute, storage, networking and GPUs, with local capacity first and cloud overflow when needed. Every command returns status, data and explicit next steps.

Infrastructure an agent can operate, inside guard rails.

Open source AI-to-CAD bridge · Python · Open source

Fusion 360 MCP

An MCP server that lets an AI assistant inspect and manipulate a live Autodesk Fusion 360 session: sketches, features, measurements, interference checks, assemblies and joints.

Conversation becomes inspectable geometry.

you › L-bracket, 40 × 40 mm, 3 mm, two M5 holes

  1. query get_design_state()
  2. create create_sketch(plane: XZ)
  3. create extrude(profile, 20 mm)
  4. create create_hole(Ø5.5, count: 2)
  5. validate measure(hole → edge)
  6. validate check_interference(M5)
  7. viewport capture_screenshot()
Run it in the live console
Open source Fabrication tooling · Open source · Beta

FlatCAM Carvera processors

FlatCAM post-processors that emit Smoothieware G-code for the Makera Carvera Air: PCB isolation milling, drilling and laser, with a two-pass auto-probe toolchange for multi-tool jobs.

Board design to machined PCB without hand-editing G-code.

2-pass auto-probe toolchange

    Also in the lab · Private R&D

    Model adaptation & evaluation

    Reproducible loops for adapting models to specific tasks: data curation, training, serving and held-out evaluation treated as one system.

    Embedded intelligence

    Sensor-rich devices and distributed control, with firmware, feedback and safety designed together from the start.

    Approach · Four invariants

    A demo is easy. Trust is engineered.

    Anyone can wire an LLM to an API. A system that behaves honestly under uncertainty, failure and real consequences takes a different set of decisions. These are ours.

    1. 01

      Name the uncertain boundary

      Stochastic work should be visible in the architecture, not disguised as a normal function call.

    2. 02

      Validate before effect

      Reasoning can explore. A typed or physical action must pass an explicit boundary first.

    3. 03

      Keep a human ceiling

      When confidence or consequence demands it, the system escalates instead of improvising.

    4. 04

      Trace what actually happened

      Decisions, costs, fallbacks and execution paths leave evidence an operator can inspect.

    Contact · Start a conversation

    Have a product idea that needs AI? Let's build it.

    Tell us what you want to automate or build. We reply with honest next steps, even if that means "you don't need an LLM for this".

    [email protected]