Gorilla

Gorilla

Open-source LLM from UC Berkeley that turns natural language into API calls and tool-using agent actions.

62/100MonitorFreeFree

If your team writes tool-calling agents and wants the model in your own infrastructure, Gorilla is the most credible open-source starting point: OpenFunctions-v2 was trained natively for parallel and multiple-function calls across Python, Java, JavaScript and REST, and BFCL gives you reproducible comparison numbers instead of vendor claims. GoEX's undo and damage-confinement abstractions are a real answer to agents taking irreversible actions. The trade is effort: you self-host, you handle latency and maintenance, and you should not expect general chat quality. Teams that want a managed path should look at OpenAI's function calling or the Vercel AI SDK instead.

Verified 5d ago · liveness 62/100 · cite: rightaichoice.com/tools/gorilla

Best for
  • Platform and infrastructure engineers building tool-calling agents
  • Teams that want an Apache 2.0 model they can fine-tune and self-host
  • Researchers benchmarking function calling with BFCL
  • Teams building domain-specific RAG pipelines with RAFT
Not ideal for
  • Product teams that need a vendor responsible for uptime and support
  • Projects without the infrastructure to self-host and maintain a model
  • General-purpose chat, writing or open-domain assistants
Visit Website

AdvancedRunning the model: minutes — the Colab notebook needs no sign-up or install. The CLI is a single pip install gorilla-cli. Getting to first production value is longer: you must stand up hosting for the 6.91B model, adapt it to your function schemas, and evaluate its accuracy before you trust it in a workflow. GoEX integration adds the work of defining which actions are reversible and how state isCLI · API · WebAPI availableVerified 5d ago
Pricing
Free
FreeFree tier3 hidden costs
Learning curve
Advanced
Running the model: minutes — the Colab notebook needs no sign-up or install. The CLI is a single pip install gorilla-cli. Getting to first production value is longer: you must stand up hosting for the 6.91B model, adapt it to your function schemas, and evaluate its accuracy before you trust it in a workflow. GoEX integration adds the work of defining which actions are reversible and how state is
Runs on
CLIAPIWeb
API available · 3 integrations
Who it's for
Platform engineerApplied researcherEngineer shipping an autonomous action
Live sentiment
Is Gorilla actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Gorilla if you need someone else to run the model for you — it is a self-hosted research release with no managed service behind it.

The 30-second take
Biggest gripe

Self-hosting is the real bill: you pay for the GPU or compute that serves the 6.91B OpenFunctions-v2 model, which a hosted API would have absorbed.

Price reality

Gorilla carries no per-call fee — it is Apache 2.0, so your cost is the compute and staff time to serve it. That inverts the usual tradeoff: managed function-calling APIs charge per token or per call but absorb the ops, while self-hosting Gorilla is cheap at sustained volume and expensive in engineering hours. It fits teams that already run model infrastructure, and it is a poor fit for teams with no GPU budget or no one to maintain a deployment.

In short

Gorilla — Open-source LLM from UC Berkeley that turns natural language into API calls and tool-using agent actions. Best for Platform and infrastructure engineers building tool-calling agents, Teams that want an Apache 2.0 model they can fine-tune and self-host, Researchers benchmarking function calling with BFCL. Free to use.

What people actually say about Gorilla — is it worth it?

We scanned public community sources for Gorilla on Jul 18, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Our own analysis of that scan says the posts were off-subject. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.

Viability Score

62/100
Monitor

How well maintained and how widely used is Gorilla? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
not measured
Traction
100
Site health
95
User sentiment
7
What the vendor publishes
20

Last calculated: October 2026

How we score →

Key Features

  • OpenFunctions-v2 model, 6.91B parameters, on HuggingFace (gorilla-llm/gorilla-openfunctions-v2)
  • Native parallel function calling — generates multiple function calls at once
  • Multiple function selection — picks one or more functions from a supplied set
  • Function relevance detection — declines when no provided function fits the question
  • Generates calls for Python, Java, JavaScript and REST APIs with extended data types
  • Berkeley Function Calling Leaderboard (BFCL) with ~2k question-function-answer pairs
  • BFCL evaluation across Python, Java, JavaScript and REST API domains
  • HuggingFace BFCL evaluation dataset and Gradio demo space
  • GoEX runtime for executing LLM-generated code and API calls
  • Undo abstraction for reverting an executed LLM action
  • Damage confinement abstraction for bounding the risk of an action
  • RAFT — Retriever-Aware FineTuning recipe for domain-specific RAG
  • Gorilla CLI, installed via pip install gorilla-cli
  • Gorilla-Spotlight search integration
  • Web demo and no-signup Colab notebook for trying the model

About Gorilla

FreeAdvancedAPI availableCLI · API · Web

Gorilla is an open-source large language model from UC Berkeley's Sky Computing Lab, built specifically for API function calling and tool use. You give it a natural-language request and a set of functions; it returns the exact call to make. Version 2 of OpenFunctions (the open-source release, 6.91B parameters, gorilla-llm/gorilla-openfunctions-v2 on HuggingFace) was natively trained for parallel function calling (multiple functions in a single response) and multiple-function selection, and added Java, REST and Python APIs for the first time with extended data types. The project is paired with the Berkeley Function Calling Leaderboard (BFCL), which benchmarks models on roughly 2,000 question-function-answer pairs across Python, Java, JavaScript and REST, covering multiple and parallel calls plus function-relevance detection when none of the supplied functions actually fit. Two companion pieces extend the stack: GoEX, an Apache 2.0 runtime for executing LLM-generated actions with undo and damage-confinement abstractions, and RAFT (Retriever-Aware FineTuning), a recipe for tuning a base model against a specific document set for RAG. Entry points include a CLI (pip install gorilla-cli), a Spotlight Search integration (Gorilla-Spotlight signup), a web demo and a Colab notebook that runs with no sign-up or install. Everything ships under Apache 2.0, so you can fine-tune, self-host and use it commercially. It is not a managed service — you run and maintain the model yourself, which suits engineering teams and researchers who want control rather than a per-call vendor.

Behind the Verdict

Gorilla's value is not that it is a chatbot — it is that it is one of the few openly licensed models built end-to-end for producing API calls, plus a benchmark suite to prove whether any model is actually good at it. OpenFunctions-v2 was trained natively for parallel functions (several calls at once) and multiple functions (choosing one or more from a supplied set), and it extended coverage to Java and REST alongside Python, with function-relevance detection so the model can decline when none of the provided functions fit the question. The Berkeley Function Calling Leaderboard turns that into an apples-to-apples number across roughly 2,000 question-function-answer pairs and multiple languages. Where Gorilla gets interesting for production work is the surrounding tooling. GoEX approaches agent safety differently from the usual guardrail stack: instead of trying to validate every intermediate step before execution, it assumes post-facto validation and gives you undo plus damage confinement, so you can revert an action or bound its blast radius after you see the result. RAFT is the third leg — a fine-tuning recipe that trains a base model against a specific document set rather than generic instruction data, intended for domain RAG where you know the corpus. Strengths: Apache 2.0 across the model, the execution engine and the reproducibility code; no per-call fee; you can fine-tune and self-host; the leaderboard and datasets are public, so you can verify claims yourself. Multiple low-friction doors in — pip install gorilla-cli, a Spotlight Search integration, a web demo and a Colab notebook with no sign-up. Weaknesses are structural, not fixable by configuration. You own deployment, latency and maintenance. The model is specialized for function calling, so it is not the right pick for open-domain conversation or general writing. GoEX's docs themselves lay out the hard parts of post-facto validation — hallucination, stochasticity, and downstream effects you cannot see while they happen — meaning undo and damage confinement are mitigations, not a solved problem. Undo also has limits: the GoEX write-up notes it depends on the level of system access and can require keeping multiple versions of state, which costs memory and compute. Where it fits: platform and infrastructure engineers building tool-calling agents, researchers who need a reproducible benchmark, and teams building domain RAG who can invest in fine-tuning. Where it does not: product teams that need a vendor on the hook for uptime, and anyone who needs a general-purpose assistant out of the box.

Researching Gorilla? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Gorilla actually fits — and what changes day-one when you adopt it.

Platform engineer

You have an internal service catalog and want an agent that turns a request like 'book the next available slot' into the correct REST call. You pull gorilla-openfunctions-v2 from HuggingFace, feed it your function schemas, and evaluate the output against a held-out set before wiring it in.

Outcome: You get the correct function call plus a measured accuracy number you can defend, using a model you host yourself.

Applied researcher

You need to argue that your agent's tool-calling accuracy holds up. You run the model through the Berkeley Function Calling Leaderboard dataset, which covers Python, Java, JavaScript and REST across multiple and parallel calls, and compare it against published scores.

Outcome: A reproducible comparison rather than a vendor's own claim, with the dataset and code available from the project.

Engineer shipping an autonomous action

Your agent will write files and call APIs without a human checking first. You wrap the generated action in the GoEX runtime, execute it, inspect the result, then either commit or undo — relying on damage confinement to bound the worst case.

Outcome: Actions become reversible instead of final, which lets you take a step toward human-out-of-the-loop execution without an unbounded blast radius.

Use Cases

Models Under the Hood

gorilla-openfunctions-v2 (6.91B parameters)

as of 2026-09-25

Limitations

  • Gorilla is specialized for function calling and API integration, so it is not built for general language tasks or open-domain conversation.
  • You are responsible for hosting it, which brings latency and infrastructure overhead a managed API would absorb.
  • The GoEX paper is candid that post-facto validation has open problems: hallucination, stochasticity, delayed feedback and downstream effects that stay invisible while the LLM works.
  • Undo is also conditional — its feasibility depends on how much system access is granted, and maintaining the state needed to revert actions costs memory and compute.
  • Performance on non-function-calling benchmarks is not covered in the documentation available to us, and the integration list here is limited to what the GoEX write-up names (Slack, Spotify, Dropbox as example bots), not a published catalog.

as of 2026-10-04

Verification history

We have re-verified Gorilla 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-checked, vendor evidence unchanged
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Self-hosting is the real bill: you pay for the GPU or compute that serves the 6.91B OpenFunctions-v2 model, which a hosted API would have absorbed.
  • Undo in GoEX can require keeping multiple versions of system state, so the safety feature trades memory and compute cost for reversibility.
  • Someone on your team owns deployment, upgrades and evaluation upkeep — that maintenance time is the operating cost, not a license fee.

Where the pricing makes sense

The company stage and team size where Gorilla's pricing actually pencils out — and where peers do it cheaper.

Gorilla carries no per-call fee — it is Apache 2.0, so your cost is the compute and staff time to serve it. That inverts the usual tradeoff: managed function-calling APIs charge per token or per call but absorb the ops, while self-hosting Gorilla is cheap at sustained volume and expensive in engineering hours. It fits teams that already run model infrastructure, and it is a poor fit for teams with no GPU budget or no one to maintain a deployment.

Setup time & first value

How long it actually takes to get something useful out of Gorilla — broken out by persona, not the marketing-page minute.

Running the model: minutes — the Colab notebook needs no sign-up or install. The CLI is a single pip install gorilla-cli. Getting to first production value is longer: you must stand up hosting for the 6.91B model, adapt it to your function schemas, and evaluate its accuracy before you trust it in a workflow. GoEX integration adds the work of defining which actions are reversible and how state is

Switching to or from Gorilla

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a hosted function-calling API like OpenAI's: port your tool schemas to Gorilla's function format and self-host the model, then run BFCL-style evaluation to check accuracy before cutting over.
  • →From prompt-only tool calling on a general LLM: fine-tune or prompt OpenFunctions-v2 with your actual function set and use function relevance detection to handle the 'no suitable function' case.
  • →From a hand-rolled RAG pipeline: follow the RAFT recipe to fine-tune the base model against your specific document set instead of relying on generic instruction tuning.
Migrating out
  • ↗To OpenAI's function calling: swap Gorilla's function-calling output for the hosted API when you want to stop maintaining model infrastructure.
  • ↗To the Vercel AI SDK: move tool-calling logic into the SDK's abstractions if you want a managed application-layer toolchain rather than a self-hosted model.

Integrations

SlackSpotifyDropbox

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Gorilla”, and we withheld 6: 6 could not be judged, because “Gorilla” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Gorilla.

Official links

Tools that pair well with Gorilla

Common stack mates teams adopt alongside Gorilla, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Gorilla

View all
Mirascope

Mirascope

Mirascope is an open-source Python library that turns LLM calls, tools, and prompt versioning into plain decorated functions.

FreemiumTry
Marvin

Marvin

Marvin is an open-source Python framework that turns ordinary functions into AI-powered tools using decorators like @ai_fn and @ai_classifier.

FreeTry
MetaGPT

MetaGPT

Open-source multi-agent framework that assigns PM, architect, engineer and QA roles to LLMs for structured software tasks

FreeTry

Frequently Asked Questions

Used Gorilla? Help shape our editorial sentiment research.