Batch inference / pilot access

Useful AI.
By the batch.

Classify, extract and enrich data with open-weight models. Start with a small batch. Measure the quality and cost before you scale.

For the work that needs to get done, not answered in a chat.

reviews.jsonl20 rows

“Great quality. Delivery took
longer than expected.”

One batch.Many small decisions.
results.jsonlstructured
topicdelivery
sentimentmixed
custom_idreview-001

01 / Illustrative input and output

Built around the job

Open-weight models Asynchronous workflows Measurable results

01 / The work

Less conversation.
More completed rows.

The same task, repeated across a dataset. That’s where batch inference belongs.

Review classificationIllustrative example

Input text

“Great quality. Delivery took longer than expected.”

Identify the topic and sentiment.
{
  "custom_id": "review-001",
  "topic": "delivery",
  "sentiment": "mixed"
}

A sample of the intended task, not a live model response. Output quality is evaluated during your pilot.

Documents, without the hand-waving. Start with extracted text. Image understanding and direct PDF workflows are on the roadmap, with model-specific evaluation before release.

02 / The interface

A batch is a job.
Not 20 browser tabs.

Prepare a JSONL file, submit it through the Python client, and collect the results. Keep a stable ID for every record so the output fits back into your pipeline.

Read the SDK guide

Local SDK preview. Requires an operator-configured profile and available model; not a hosted API or public package install.

batch.py
# Configured pilot environment
from firmbatch import LocalClient

client = LocalClient()
job = client.batches.create(
    model="llama-3.1-8b-instruct",
    input_file="reviews.jsonl",
)
job.wait()
if job.output_file is not None:
    job.download("results.jsonl")
SDK preview Setup and limitations in the guide ↗

03 / Location matters

Swiss requirements.
Specific answers.

For teams with Swiss data-location requirements, we’re developing a Swiss-hosted inference option. The pilot starts with your requirements, not a blanket compliance claim.

  • Agree the location of inference and data storage.
  • Map logs, backups, access and subprocessors.
  • Evaluate the model on your acceptance criteria.

Infrastructure is still being set up. Swiss GPU processing alone does not make the full service Swiss-only.

Discuss a Swiss pilot

Need a different location?
Tell us where the workload is allowed to run.

Switzerland Pilot setupEuropean Union PlannedUnited States PlannedAsia Planned

This is the intended regional offering, not a live endpoint catalogue. Countries, models and capacity must be confirmed for each pilot.

04 / Prove the economics

The price that matters?
Cost per useful result.

Tokens are an input. Correctly labelled reviews, extracted fields and usable records are the outcome. A pilot gives us the evidence to price your actual workload.

Get a scoped pilot proposal
  1. 01

    Bring a representative sample

    Start with 20 records and a clear definition of a good result. Use non-sensitive or synthetic data first.

  2. 02

    Measure quality, time and cost

    Agree the model and region. Compare the outputs against your existing process.

  3. 03

    Decide whether to scale

    Agree scope and commercial terms before a larger run. No published token rate or SLA yet.

05 / A focused start

One useful product first.

Batch inference is the starting point. The rest has to earn its place.

Pilot focus

Batch text inference

Classification, extraction and enrichment with open-weight models.

Planned

Decision & retrieval models

Decision models such as Clef, plus embeddings and reranking. Separate integrations, not available endpoints today.

Planned

Multimodal & fine-tuning

Image and document workflows, then task-specific model adaptation. Availability will follow evaluation.

Looking for async containers or code review? Those are exploratory ideas, not part of the current offer.

The practical questions

Before you batch.

Can I sign up and run jobs today?

Access is currently through an operator-led pilot. The batch runner exists, but hosted API access, account onboarding and billing are not yet a public self-service product. We’ll agree setup and scope with you first.

Is this an OpenAI-compatible Batch API?

Not a drop-in replacement. The current JSONL input uses message-based request bodies, but the job lifecycle belongs to Firmbatch’s own Python client. The developer guide documents that distinction.

Which model should I use?

That depends on the task, the quality threshold and where your data can run. We start with an agreed open-weight text model and compare results on your sample. Model names on the roadmap are not a guarantee of available capacity.

Is the Swiss option ready for regulated data?

Not yet. The infrastructure is in development. A regulated deployment needs a review of the full data flow, contracts and operational controls, not just the GPU location. Don’t send sensitive data in a pilot enquiry.

Are you promising the cheapest tokens?

No. GPU price alone doesn’t determine your total cost. Model quality, batching, startup time, retries and utilisation all matter. We want to compare the cost of results you can actually use.

Your data. A concrete next step.

Start small.
Make it useful.

Tell us what you need to process, roughly how much, and where the data is allowed to go. We’ll work out a pilot from there.

Let’s talk about your batch

An email, not a sales funnel. No sensitive data or attachments needed.