Insights

Private inference vs. public model APIs, explained for healthcare

Two AI tools can look identical and handle data very differently. Here is what changes when the model runs privately, and how to tell.

Put two AI assistants side by side and they can look identical: a text box, a model that answers well, a place to upload documents. For a healthcare team, the most important difference between them is invisible on screen. It is where the model actually runs, and therefore where your information goes when someone presses Enter.

This piece explains the two common designs in plain terms, what each means for protected health information, and the questions that reveal which one a vendor is really offering.

Most AI products don’t run their own model

Large language models are expensive to build and operate, so many AI products don’t host one. Instead, the product sends each request to a model provider through an API, an interface that accepts text and returns a response, and displays the result.

With this design, every prompt leaves the product you are using. The text a user typed, and any document they attached, travels to the model provider’s systems, is processed there, and comes back. The product you evaluated is one company; the model that read your data may belong to another.

What a public model API means for PHI

Sending PHI through a third-party model API is not automatically noncompliant. Some model providers offer Business Associate Agreements and specific data-handling commitments for particular services. But the design has consequences that your privacy and security teams will want to understand:

  • Another party handles the data. The model provider becomes a subcontractor in the chain, and HIPAA requires the product vendor to obtain written assurances from subcontractors that handle PHI. You are now relying on two sets of commitments, not one.
  • Retention may happen somewhere you don’t see. Providers can keep requests for purposes such as abuse monitoring, under their own terms. Those terms are part of your risk, even though you never signed them.
  • The infrastructure is shared. Public model services are, by design, used by many customers at once. Separation between customers is the provider’s responsibility and happens inside systems you cannot inspect.
  • The answer can change. Which provider a product uses, and under what terms, is a business decision the vendor can revisit.

None of this makes the design wrong. It makes it harder to explain, and harder to verify.

What private inference means

With private inference, the model runs inside the environment that serves your organization. Prompts and documents are processed there and are not sent out to a public third-party model service.

The practical effect is a shorter, more contained path for your data:

  • fewer companies handle PHI
  • the systems that read your information sit inside the same boundary as your stored data and access controls
  • the answer to “where did that prompt go?” fits in one sentence

For a security review, that simplicity matters. A data flow you can draw on one page is a data flow you can reason about.

The trade-offs, honestly

Private inference is not free of compromises, and a vendor who says otherwise is overselling.

  • Model choice. Public APIs give the fastest access to the newest, largest models. Private deployments choose models more deliberately, and a new release may not be available the day it is announced.
  • Operations. Someone has to run, patch, monitor, and secure the infrastructure the model runs on. With private inference, that responsibility sits with the vendor, so ask how they carry it.
  • It doesn’t replace the basics. Where the model runs is one control among many. Access management, retention, encryption, and human review of outputs matter just as much, whichever design you choose.

Questions that reveal which one you’re buying

Architecture diagrams in sales decks tend to be reassuring and abstract. These questions get to the real answer:

  1. When a user submits a prompt, which systems does the text pass through, and which companies operate them?
  2. Is any customer data sent to a third-party model provider? If so, which one, and under what agreement?
  3. Where are prompts and outputs logged, and for how long?
  4. Is the model shared with other customers, and can anything we submit influence it?
  5. Can you show us the data flow for a single request, end to end?

A vendor using private inference will answer these quickly. A vendor using public APIs can still answer them well, and if they can’t, that tells you something too.

Where Timberline fits

Timberline uses private inference. Each customer’s work runs in a logically isolated Timberline environment, users sign in through their organization’s SSO, and data is encrypted with keys dedicated to that customer. Production PHI is not sent to a public third-party hosted model API, and customer information is never used to train models.

We chose this design because it is the one we would want to defend in our own security review. See how the architecture fits together, or review the full control inventory in the Trust Center.

This article is general information, not legal advice. Review specific agreements and deployments with your own counsel and compliance team.

  1. 5 min read

    What to ask before your team uses AI with PHI

    Six questions that show whether an AI tool can safely handle protected health information, and what good answers sound like.

  2. 4 min read

    Three months of building Timberline

    From a private AI workspace in July to checked citations, projects, and governed web search in September: what we built, and why.

See Timberline in practice.

Talk with Timberline