The LLM Data Company

The LLM Data Company

Paper Instruments trains open-source, domain-specific frontier models and agent tooling for specialist knowledge work.

54/100UnverifiedCustom pricingContact Sales

A research-first lab with a real medical model on the board and open, inspectable methods before you commit budget. Kos-1 Lite and the Kos-1 Experimental 1T-parameter RL result on Kimi K2.5 are genuine receipts, and DiligenceBench plus DRACO give evaluation teams something to run. The catch is timing: Feather, Ultramarine, and the Paper Instruments rebrand are all marked coming soon, so most of the buyer-facing stack is not shippable today. Pick it if you have ML resources and a production harness already; otherwise wait for the launch and keep a generalist provider as your safety net.

Last checked 10d ago · cite: rightaichoice.com/tools/the-llm-data-company

Best for
  • Enterprise teams running production agents in healthcare or finance
  • Organizations with an existing production harness wanting custom specialist models
  • Teams building verifiable domain agents in math or code that need high accuracy
  • Researchers evaluating long-form equity-research agents with DiligenceBench
Not ideal for
  • Individual developers or small teams without a production harness or ML staff
  • Buyers who need a shippable agent harness today — Feather and Ultramarine are coming soon
  • Projects requiring low-effort deployment with no custom training or evaluation work
Visit Website

AdvancedFor an ML team with an existing harness: a few days to evaluate Kos-1 Lite against your own test set and stand up the published rubric judge. For teams waiting on Feather or Ultramarine: no ETA, both are marked coming soon. For Paper Office, an afternoon of Python integration for an engineer comfortable with agent libraries.API · WebAPI availableLast checked 10d ago
Pricing
Custom pricing
Contact Sales3 hidden costs
Learning curve
Advanced
For an ML team with an existing harness: a few days to evaluate Kos-1 Lite against your own test set and stand up the published rubric judge. For teams waiting on Feather or Ultramarine: no ETA, both are marked coming soon. For Paper Office, an afternoon of Python integration for an engineer comfortable with agent libraries.
Runs on
APIWeb
API available
Who it's for
ML lead at a hospital systemQuant researcher at an equity research deskEngineer building document automation
Live sentiment
Is The LLM Data Company actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Paper Instruments if you need a ready-to-deploy agent harness or model line this quarter — Feather and Ultramarine are still marked coming soon and the stack is mid-rebrand.

The 30-second take
Biggest gripe

Domain training runs consume significant compute, and the Kos-1 Experimental work scaled to 1 trillion parameters — your own training bill will reflect that class of run.

Price reality

Pricing is handled through direct vendor contact, so this does not slot into a per-seat comparison against ChatGPT Team or Claude Team. Compare it instead against the GPU and evaluation spend of training your own domain model in-house, or against the premium an OpenAI or Anthropic deployment charges for a narrow high-accuracy workload. Budget for services-style engagement rather than a monthly subscription line item.

In short

The LLM Data Company — Paper Instruments trains open-source, domain-specific frontier models and agent tooling for specialist knowledge work. Best for Enterprise teams running production agents in healthcare or finance, Organizations with an existing production harness wanting custom specialist models, Teams building verifiable domain agents in math or code that need high accuracy. Contact Sales pricing.

What people actually say about The LLM Data Company — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

19 mentions across 2 sources (YouTube, Lemmy) · researched Sep 24, 2026.

42% positive58% critical

Weighted by the 27 posts each of 2 sources contributed.

Recurring strengths
  • +Open-source model releases let enterprises inspect and verify claims internally rather than trust marketing
  • +Domain focus on healthcare and finance targets regulated environments where generalist models underperform
  • +Kos-1 Experimental reportedly scaled environment-free RL to 1T parameters — a genuine efficiency claim
  • +Curriculum autoresearch system curates tasks and rewards, potentially reducing hallucination risk in niche work
  • +DiligenceBench and DRACO with Perplexity add transparency to how long-form research quality is measured
Recurring frustrations
  • −No real community reviews exist — the scraped posts are all keyword coincidences, not product feedback
  • −Feather and Inkwell aren't released yet, so the marketed product line is largely speculative
  • −Pricing is 'contact us' only, with no published tiers, trial, or transparent cost structure
  • −Benchmark credibility leans on metrics the company itself authors or co-publishes
  • −Listed as 'advanced' with no documented integrations, implying meaningful custom engineering
Patterns worth knowing
Scraped posts are keyword coincidences, not actual discussion of this tool
Seen on YouTube, Lemmy
Rising concern about AI agents executing untrusted code in corporate environments
Seen on Lemmy
Demand for practical, hands-on LLM engineering content over theoretical framing
Seen on YouTube
Learning curve
advancedProductive in ~Days of setup
Hidden costs people mention
  • • No self-serve tier — evaluation requires a sales conversation, adding procurement time
  • • Likely integration/engineering cost since no off-the-shelf integrations are listed
  • • Potential compute costs to run and fine-tune models in your own harness
  • • Feather and Inkwell may carry separate pricing once released

Viability Score

54/100
Unverified

How well maintained and how widely used is The LLM Data Company? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
40
identity move
not measured
User sentiment
43
What the vendor publishes
20

Last calculated: October 2026

How we score →

Key Features

  • Open-source frontier models for domain-specific knowledge work
  • Kos-1 Lite medical model for healthcare AI
  • Kos-1 Experimental: 1T-parameter RL training on Kimi K2.5
  • On-policy reinforcement learning for domain specialization
  • Environment-free reinforcement learning at scale
  • Paper Office: agent-first Python library for document creation and editing
  • Feather: agent harness built for knowledge work (coming soon)
  • Ultramarine: frontier models for knowledge work (coming soon)
  • Curriculum autoresearch system for task and reward curation
  • DiligenceBench: benchmark for long-form equity-research agents
  • DRACO: deep research evaluation benchmark built with Perplexity
  • Published rubric judge training methodology
  • Open research notes, methods, and results
  • Reduced serving cost vs generalist frontier models

About The LLM Data Company

Contact SalesAdvancedAPI availableAPI · Web

Paper Instruments (formerly The LLM Data Company) is a research-first lab that trains open-source frontier models for knowledge work instead of optimizing for generalist benchmarks. Its approach is narrow by design: pick a domain such as healthcare or finance, train inside a real production harness using on-policy reinforcement learning, and let accuracy and verifiability outweigh breadth. The public product surface currently splits into Paper Office, an agent-first Python library for programmatic document creation and editing, plus three tracks listed as coming soon on the vendor site — Feather (an agent harness built for knowledge work), Ultramarine (frontier models for knowledge work), and the company's rebrand to Paper Instruments. Research is published openly: DiligenceBench evaluates long-form equity-research agents, DRACO was built with Perplexity to target deep research evaluation, and rubric judge training notes are published for teams building their own evaluation pipelines. The concrete model receipts are Kos-1 Lite, shipped as a leading medical model, and Kos-1 Experimental, which scaled environment-free reinforcement learning to 1 trillion parameters on Kimi K2.5. Training tasks and rewards come from Curriculum, the company's autoresearch system. It suits teams that already run a production harness and want cheaper, more accurate domain models than a generalist frontier provider delivers.

Behind the Verdict

Paper Instruments is best understood as a specialist model lab rather than a product company you sign up for and use on Monday. The strongest evidence that the specialization claim is real is Kos-1 Lite, shipped as a leading medical model, and Kos-1 Experimental, which scaled environment-free reinforcement learning to 1 trillion parameters on Kimi K2.5 — a training-efficiency result that matters if your budget is the constraint. Curriculum, the company's autoresearch system, is the mechanism that generates tasks and rewards, and it's the piece you'd actually inherit if you bought into the domain-training pitch. Where it differs from generalist providers is verifiability: the published rubric judge training notes, DiligenceBench for long-form equity-research agents, and DRACO (built with Perplexity) for deep research evaluation give an internal eval team real scaffolding instead of a vendor benchmark you have to take on faith. The weaknesses are real too. Feather, the agent harness for knowledge work, and Ultramarine, the frontier model line, both carry a coming-soon label, and the whole company is mid-rebrand to Paper Instruments, so documentation and access paths are in transition. The only currently described shipping artifact is Paper Office, an agent-first Python library for document creation and editing, which implies an engineering audience comfortable writing code. If you don't run your own evals, most of this lab's value is unclaimable — its whole argument is accuracy you can prove internally. Fit it into a stack where a generalist model handles the breadth and a Kos-class specialist handles the narrow, high-accuracy slice, and it reads as a credible shortlist candidate. Expect to contact the vendor directly for access terms rather than self-serve through a pricing page.

Researching The LLM Data Company? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas The LLM Data Company actually fits — and what changes day-one when you adopt it.

ML lead at a hospital system

You already run a clinical documentation harness and want a medical specialist instead of a generalist model. You pull Kos-1 Lite, wire it into the harness, and run your own test set against the published rubric judge methodology.

Outcome: A domain model you can benchmark internally before it touches patient-facing workflows, at lower serving cost than the generalist model it replaces.

Quant researcher at an equity research desk

You need to know whether a long-form research agent's output is trustworthy. You run the agent against DiligenceBench and use the published rubric judge notes to score it.

Outcome: A defensible, repeatable evaluation of your research agent rather than an anecdotal read of its summaries.

Engineer building document automation

You install Paper Office, the agent-first Python library, and script programmatic creation and editing of office documents inside your existing pipeline.

Outcome: Document generation that runs in code and stays inside your own harness, with an upgrade path to Feather when it ships.

Use Cases

Models Under the Hood

Kimi K2.5Kos-1 LiteKos-1 Experimental

as of 2026-10-08

Limitations

  • The current site is largely a holding page: Ultramarine, Feather, and Paper Office-adjacent tooling are listed as coming soon or in early announce-stage, so buyer-facing products are incomplete.
  • Kos-1 Lite and Kos-1 Experimental are described as state-of-the-art or experimental, so production stability is not guaranteed.
  • Paper Office is delivered as Python packages/libraries for Word, PowerPoint, and Excel, implying engineering work before value.
  • The evidence shows no self-serve pricing or signup path, so onboarding appears to depend on contacting the vendor.

as of 2026-09-28

Verification history

We have re-verified The LLM Data Company 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 7 verification passes.

Free to cite with attribution — this page re-verifies continuously.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Domain training runs consume significant compute, and the Kos-1 Experimental work scaled to 1 trillion parameters — your own training bill will reflect that class of run.
  • Adapter and eval work falls on your team: Paper Office is a Python library and the published rubrics assume you can staff an internal evaluation pipeline.
  • Running Kos-1 Lite alongside a generalist model for breadth means paying two serving bills during the transition period.

Where the pricing makes sense

The company stage and team size where The LLM Data Company's pricing actually pencils out — and where peers do it cheaper.

Pricing is handled through direct vendor contact, so this does not slot into a per-seat comparison against ChatGPT Team or Claude Team. Compare it instead against the GPU and evaluation spend of training your own domain model in-house, or against the premium an OpenAI or Anthropic deployment charges for a narrow high-accuracy workload. Budget for services-style engagement rather than a monthly subscription line item.

Setup time & first value

How long it actually takes to get something useful out of The LLM Data Company — broken out by persona, not the marketing-page minute.

For an ML team with an existing harness: a few days to evaluate Kos-1 Lite against your own test set and stand up the published rubric judge. For teams waiting on Feather or Ultramarine: no ETA, both are marked coming soon. For Paper Office, an afternoon of Python integration for an engineer comfortable with agent libraries.

Switching to or from The LLM Data Company

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a generalist frontier API: route your narrow high-accuracy tasks to a Kos-class specialist while keeping the generalist for breadth.
  • →From an in-house domain model: adopt Curriculum-generated tasks and rewards instead of hand-building your own training set.
  • →From ad-hoc agent scoring: replace gut-feel review with DiligenceBench or a rubric judge built from the published training notes.
Migrating out
  • ↗To OpenAI or Anthropic: move back to a generalist API if you need many task types deployable immediately.
  • ↗To a managed agent platform: switch if you want an agent harness you don't have to build or wait for.

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “The LLM Data Company”, and we withheld 6: 6 did not mention The LLM Data Company. We are showing none, because we could not prove any of them are about The LLM Data Company.

Tools that pair well with The LLM Data Company

Common stack mates teams adopt alongside The LLM Data Company, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

The Llm Data Company vs Spider Cloud

Choose Spider Cloud if you need affordable, high-speed web data for AI agents or RAG—its freemium pricing and 1,000+ scrapers are unmatched. Choose The LLM Data Company if you're an enterprise in healthcare or finance needing a custom-trained specialist model that beats GPT-4/Claude at lower cost; their Kos-1 Lite and Experimental models prove their approach. These tools address entirely different needs—data access vs. model specialization—so your decision hinges on whether your bottleneck is gathering data or training a domain-specific model.

The Llm Data Company vs Temporal Ai

For teams building reliable AI agents that survive crashes and require orchestration, Temporal is the clear choice—its open-source durability and workflow capabilities are unmatched. If your priority is domain-specific model specialization (e.g., medical reasoning) and you have a production harness, The LLM Data Company offers cutting-edge training that can outperform generalist models at lower cost. Most buyers will start with Temporal for orchestration and only consider The LLM Data Company for niche, high-stakes domain specialization.

The Llm Data Company vs Presto Voice

Choose Presto Voice if you run a QSR drive-thru chain and need a drop-in voice AI that boosts order accuracy and upsell revenue. Choose The LLM Data Company if you need a specialized, cost-effective model for a critical domain like healthcare, where outperforming GPT-4o is a priority. The two tools serve entirely different markets—restaurants vs. enterprise AI agents—so your choice depends on whether your problem is at the drive-thru window or in the production ML pipeline.

Alternatives to The LLM Data Company

View all
AfterQuery

AfterQuery

Applied research lab that captures expert reasoning and structures it into SFT, RL rubric, agent, and computer-use training data for frontier models.

Contact SalesTry
Magic.dev

Magic.dev

Frontier code models built for ultra-long-context software engineering and AI research automation.

Contact SalesTry
Falcon LLM

Falcon LLM

Apache 2.0 open-weight model family from TII Abu Dhabi, spanning hybrid Transformer-Mamba, Arabic, reasoning, and multimodal vision models.

FreeTry

Used The LLM Data Company? Help shape our editorial sentiment research.