Gorilla
Open-source LLM from UC Berkeley that turns natural language into API calls and tool-using agent actions.
If your team writes tool-calling agents and wants the model in your own infrastructure, Gorilla is the most credible open-source starting point: OpenFunctions-v2 was trained natively for parallel and multiple-function calls across Python, Java, JavaScript and REST, and BFCL gives you reproducible comparison numbers instead of vendor claims. GoEX's undo and damage-confinement abstractions are a real answer to agents taking irreversible actions. The trade is effort: you self-host, you handle latency and maintenance, and you should not expect general chat quality. Teams that want a managed path should look at OpenAI's function calling or the Vercel AI SDK instead.
Verified 5d ago · liveness 62/100 · cite: rightaichoice.com/tools/gorilla
- Platform and infrastructure engineers building tool-calling agents
- Teams that want an Apache 2.0 model they can fine-tune and self-host
- Researchers benchmarking function calling with BFCL
- Teams building domain-specific RAG pipelines with RAFT
- Product teams that need a vendor responsible for uptime and support
- Projects without the infrastructure to self-host and maintain a model
- General-purpose chat, writing or open-domain assistants
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Gorilla if you need someone else to run the model for you — it is a self-hosted research release with no managed service behind it.
Self-hosting is the real bill: you pay for the GPU or compute that serves the 6.91B OpenFunctions-v2 model, which a hosted API would have absorbed.
Gorilla carries no per-call fee — it is Apache 2.0, so your cost is the compute and staff time to serve it. That inverts the usual tradeoff: managed function-calling APIs charge per token or per call but absorb the ops, while self-hosting Gorilla is cheap at sustained volume and expensive in engineering hours. It fits teams that already run model infrastructure, and it is a poor fit for teams with no GPU budget or no one to maintain a deployment.
In short
Gorilla — Open-source LLM from UC Berkeley that turns natural language into API calls and tool-using agent actions. Best for Platform and infrastructure engineers building tool-calling agents, Teams that want an Apache 2.0 model they can fine-tune and self-host, Researchers benchmarking function calling with BFCL. Free to use.
What people actually say about Gorilla — is it worth it?
We scanned public community sources for Gorilla on Jul 18, 2026 and could not establish that the discussion we found is about this tool rather than something else sharing its name. Our own analysis of that scan says the posts were off-subject. Rather than publish a sentiment score built on the wrong subject, we publish nothing here and re-run the scan.
Viability Score
How well maintained and how widely used is Gorilla? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- OpenFunctions-v2 model, 6.91B parameters, on HuggingFace (gorilla-llm/gorilla-openfunctions-v2)
- Native parallel function calling — generates multiple function calls at once
- Multiple function selection — picks one or more functions from a supplied set
- Function relevance detection — declines when no provided function fits the question
- Generates calls for Python, Java, JavaScript and REST APIs with extended data types
- Berkeley Function Calling Leaderboard (BFCL) with ~2k question-function-answer pairs
- BFCL evaluation across Python, Java, JavaScript and REST API domains
- HuggingFace BFCL evaluation dataset and Gradio demo space
- GoEX runtime for executing LLM-generated code and API calls
- Undo abstraction for reverting an executed LLM action
- Damage confinement abstraction for bounding the risk of an action
- RAFT — Retriever-Aware FineTuning recipe for domain-specific RAG
- Gorilla CLI, installed via pip install gorilla-cli
- Gorilla-Spotlight search integration
- Web demo and no-signup Colab notebook for trying the model
About Gorilla
Gorilla is an open-source large language model from UC Berkeley's Sky Computing Lab, built specifically for API function calling and tool use. You give it a natural-language request and a set of functions; it returns the exact call to make. Version 2 of OpenFunctions (the open-source release, 6.91B parameters, gorilla-llm/gorilla-openfunctions-v2 on HuggingFace) was natively trained for parallel function calling (multiple functions in a single response) and multiple-function selection, and added Java, REST and Python APIs for the first time with extended data types. The project is paired with the Berkeley Function Calling Leaderboard (BFCL), which benchmarks models on roughly 2,000 question-function-answer pairs across Python, Java, JavaScript and REST, covering multiple and parallel calls plus function-relevance detection when none of the supplied functions actually fit. Two companion pieces extend the stack: GoEX, an Apache 2.0 runtime for executing LLM-generated actions with undo and damage-confinement abstractions, and RAFT (Retriever-Aware FineTuning), a recipe for tuning a base model against a specific document set for RAG. Entry points include a CLI (pip install gorilla-cli), a Spotlight Search integration (Gorilla-Spotlight signup), a web demo and a Colab notebook that runs with no sign-up or install. Everything ships under Apache 2.0, so you can fine-tune, self-host and use it commercially. It is not a managed service — you run and maintain the model yourself, which suits engineering teams and researchers who want control rather than a per-call vendor.
Behind the Verdict
Gorilla's value is not that it is a chatbot — it is that it is one of the few openly licensed models built end-to-end for producing API calls, plus a benchmark suite to prove whether any model is actually good at it. OpenFunctions-v2 was trained natively for parallel functions (several calls at once) and multiple functions (choosing one or more from a supplied set), and it extended coverage to Java and REST alongside Python, with function-relevance detection so the model can decline when none of the provided functions fit the question. The Berkeley Function Calling Leaderboard turns that into an apples-to-apples number across roughly 2,000 question-function-answer pairs and multiple languages. Where Gorilla gets interesting for production work is the surrounding tooling. GoEX approaches agent safety differently from the usual guardrail stack: instead of trying to validate every intermediate step before execution, it assumes post-facto validation and gives you undo plus damage confinement, so you can revert an action or bound its blast radius after you see the result. RAFT is the third leg — a fine-tuning recipe that trains a base model against a specific document set rather than generic instruction data, intended for domain RAG where you know the corpus. Strengths: Apache 2.0 across the model, the execution engine and the reproducibility code; no per-call fee; you can fine-tune and self-host; the leaderboard and datasets are public, so you can verify claims yourself. Multiple low-friction doors in — pip install gorilla-cli, a Spotlight Search integration, a web demo and a Colab notebook with no sign-up. Weaknesses are structural, not fixable by configuration. You own deployment, latency and maintenance. The model is specialized for function calling, so it is not the right pick for open-domain conversation or general writing. GoEX's docs themselves lay out the hard parts of post-facto validation — hallucination, stochasticity, and downstream effects you cannot see while they happen — meaning undo and damage confinement are mitigations, not a solved problem. Undo also has limits: the GoEX write-up notes it depends on the level of system access and can require keeping multiple versions of state, which costs memory and compute. Where it fits: platform and infrastructure engineers building tool-calling agents, researchers who need a reproducible benchmark, and teams building domain RAG who can invest in fine-tuning. Where it does not: product teams that need a vendor on the hook for uptime, and anyone who needs a general-purpose assistant out of the box.
Researching Gorilla? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Gorilla actually fits — and what changes day-one when you adopt it.
You have an internal service catalog and want an agent that turns a request like 'book the next available slot' into the correct REST call. You pull gorilla-openfunctions-v2 from HuggingFace, feed it your function schemas, and evaluate the output against a held-out set before wiring it in.
Outcome: You get the correct function call plus a measured accuracy number you can defend, using a model you host yourself.
You need to argue that your agent's tool-calling accuracy holds up. You run the model through the Berkeley Function Calling Leaderboard dataset, which covers Python, Java, JavaScript and REST across multiple and parallel calls, and compare it against published scores.
Outcome: A reproducible comparison rather than a vendor's own claim, with the dataset and code available from the project.
Your agent will write files and call APIs without a human checking first. You wrap the generated action in the GoEX runtime, execute it, inspect the result, then either commit or undo — relying on damage confinement to bound the worst case.
Outcome: Actions become reversible instead of final, which lets you take a step toward human-out-of-the-loop execution without an unbounded blast radius.
Use Cases
- Generate the correct API call from a natural-language request across Python, Java, JavaScript or REST
- Build an agent that selects from a catalog of functions and calls several in parallel
- Benchmark your model's function calling against others using BFCL
- Execute LLM-generated code or API calls with GoEX and revert the action if the output is wrong
- Bound the blast radius of an autonomous agent action with damage confinement
- Fine-tune a base model against a specific document set for domain RAG with RAFT
- Run Gorilla from the terminal via the gorilla-cli package
- Try the model in a Colab notebook with no sign-up or install
Models Under the Hood
as of 2026-09-25
Limitations
- Gorilla is specialized for function calling and API integration, so it is not built for general language tasks or open-domain conversation.
- You are responsible for hosting it, which brings latency and infrastructure overhead a managed API would absorb.
- The GoEX paper is candid that post-facto validation has open problems: hallucination, stochasticity, delayed feedback and downstream effects that stay invisible while the LLM works.
- Undo is also conditional — its feasibility depends on how much system access is granted, and maintaining the state needed to revert actions costs memory and compute.
- Performance on non-function-calling benchmarks is not covered in the documentation available to us, and the integration list here is limited to what the GoEX write-up names (Slack, Spotify, Dropbox as example bots), not a published catalog.
as of 2026-10-04
Verification history
We have re-verified Gorilla 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 9 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Where the pricing makes sense
The company stage and team size where Gorilla's pricing actually pencils out — and where peers do it cheaper.
Gorilla carries no per-call fee — it is Apache 2.0, so your cost is the compute and staff time to serve it. That inverts the usual tradeoff: managed function-calling APIs charge per token or per call but absorb the ops, while self-hosting Gorilla is cheap at sustained volume and expensive in engineering hours. It fits teams that already run model infrastructure, and it is a poor fit for teams with no GPU budget or no one to maintain a deployment.
Setup time & first value
How long it actually takes to get something useful out of Gorilla — broken out by persona, not the marketing-page minute.
Running the model: minutes — the Colab notebook needs no sign-up or install. The CLI is a single pip install gorilla-cli. Getting to first production value is longer: you must stand up hosting for the 6.91B model, adapt it to your function schemas, and evaluate its accuracy before you trust it in a workflow. GoEX integration adds the work of defining which actions are reversible and how state is
Switching to or from Gorilla
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From a hosted function-calling API like OpenAI's: port your tool schemas to Gorilla's function format and self-host the model, then run BFCL-style evaluation to check accuracy before cutting over.
- →From prompt-only tool calling on a general LLM: fine-tune or prompt OpenFunctions-v2 with your actual function set and use function relevance detection to handle the 'no suitable function' case.
- →From a hand-rolled RAG pipeline: follow the RAFT recipe to fine-tune the base model against your specific document set instead of relying on generic instruction tuning.
- ↗To OpenAI's function calling: swap Gorilla's function-calling output for the hosted API when you want to stop maintaining model infrastructure.
- ↗To the Vercel AI SDK: move tool-calling logic into the SDK's abstractions if you want a managed application-layer toolchain rather than a self-hosted model.
Integrations
Resources & Guides
- Resourcegithub.com
Gorilla · Gorilla
Helpful link from github.com
- Resourcegithub.com
Goex · Gorilla
Helpful link from github.com
- Resourcehuggingface.co
Gorilla Openfunctions V2 · Gorilla
Helpful link from huggingface.co
- Resourcegorilla.cs.berkeley.edu
10 Gorilla Exec Engine · Gorilla
Helpful link from gorilla.cs.berkeley.edu
Tutorials & Learning
YouTube returned 6 videos for “Gorilla”, and we withheld 6: 6 could not be judged, because “Gorilla” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Gorilla.
Official links
Tools that pair well with Gorilla
Common stack mates teams adopt alongside Gorilla, with the specific reason each pairing earns its keep.
Mirascope
Mirascope is an open-source Python library that turns LLM calls, tools, and prompt versioning into plain decorated functions.
Marvin
Marvin is an open-source Python framework that turns ordinary functions into AI-powered tools using decorators like @ai_fn and @ai_classifier.
MetaGPT
Open-source multi-agent framework that assigns PM, architect, engineer and QA roles to LLMs for structured software tasks
Featured Head-to-Head Comparisons
Gorilla vs Locus Robotics
Locus Robotics and Gorilla serve completely different domains. Locus Robotics is for warehouses needing physical automation with a Robots-as-a-Service subscription, while Gorilla is a free open-source LLM for developers building API-calling agents. There is no direct competition; choose based on whether your problem is operational logistics or software automation.
Gorilla vs Presto Voice
Presto Voice and Gorilla serve entirely different needs. Presto Voice is a turnkey voice AI for QSR drive-thrus, offering measurable revenue lift but with enterprise pricing and limited flexibility. Gorilla is an open-source LLM for developers to add function calling to their own agents, free but requiring technical expertise. Choose based on your domain: restaurant operations or LLM application development.
Gorilla vs Truleo
Truleo and Gorilla serve completely different buyers. Truleo is a specialized, paid intelligence platform for law enforcement that automates lead generation from siloed data. Gorilla is a free, open-source LLM for developers focused on function calling and API integration. Your choice depends entirely on your domain: law enforcement or software development.
Alternatives to Gorilla
View allMirascope
Mirascope is an open-source Python library that turns LLM calls, tools, and prompt versioning into plain decorated functions.
Frequently Asked Questions
Used Gorilla? Help shape our editorial sentiment research.