Hanlp Lucene Plugin
Free HanLP Chinese word segmentation plugin for Apache Solr 5.x and Lucene 5.x full-text retrieval
If your search stack is Solr 5.x and Chinese text matters, this plugin is the least-effort route to HanLP-grade tokenization — two JARs and a schema edit. The catch is the release date: the vendor page and tutorial are 2015-era, tied to Solr 5.2.1, so newer Solr branches are unproven territory. Pick it deliberately for legacy 5.x deployments, not as a forward-looking bet.
Verified 6d ago · liveness 55/100 · cite: rightaichoice.com/tools/hanlp-lucene-plugin
- Search engineers maintaining Apache Solr 5.x with Chinese-language content
- Java developers who want HanLP tokenization configured purely through schema.xml
- Teams needing domain-specific custom dictionaries for better Chinese recall
- E-commerce or catalog search where ambiguous Chinese word boundaries cause false hits
- Deployments on Solr versions newer than 5.x without a compatibility spike first
- Non-Lucene search platforms such as Elasticsearch that need a separate adapter
- Teams wanting a managed or hosted Chinese tokenization service
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Hanlp Lucene Plugin if you are not running Solr 5.x or Lucene 5.x, or if you expect ongoing support and updates for newer versions; it's a 2015-era plugin with no recent releases.
You may need to build the plugin from source if you want to use it with newer Solr versions, which requires Java and build tooling expertise.
This plugin is free, making it cost-attractive for small teams already invested in Solr 5.x. However, it lacks the ongoing development of paid alternatives like Elastic's official Chinese analyzer or commercial NLP APIs. For budget-conscious teams on legacy Solr, it's a good fit; for those needing modern support, consider paid options.
In short
Hanlp Lucene Plugin — Free HanLP Chinese word segmentation plugin for Apache Solr 5.x and Lucene 5.x full-text retrieval. Best for Search engineers maintaining Apache Solr 5.x with Chinese-language content, Java developers who want HanLP tokenization configured purely through schema.xml, Teams needing domain-specific custom dictionaries for better Chinese recall. Free to use.
What people actually say about Hanlp Lucene Plugin — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
8 mentions across 1 source (GitHub) · researched Jul 5, 2026.
Average across the 1 source that answered — each source counts once, not each post.
- +Free and open-source with no licensing costs.
- +Integrates tightly with Lucene/Solr without code changes.
- +Supports multiple segmentation modes including NLP with POS tagging.
- +Custom dictionary support for domain-specific terminology.
- +Named entity recognition for persons, locations, and organizations.
- −Offset errors plague document ingestion from Tika.
- −Traditional Chinese tokenization is broken out of the box.
- −No built-in simplified-traditional conversion for search.
- −Documentation for custom dictionary configuration is sparse.
- −Community support with slow response and unresolved issues.
- • Time cost for debugging offset and compatibility issues
- • Need to manually add Traditional Chinese dictionaries
Viability Score
How well maintained and how widely used is Hanlp Lucene Plugin? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Chinese word segmentation with HanLP dictionary- and model-based tokenizer
- Schema-driven configuration via com.hankcs.lucene.HanLPTokenizerFactory
- indexMode expands long words into all contained sub-words at index time
- Query analyzer mode kept off to preserve PhraseQuery behavior
- Custom dictionary and configuration support through hanlp.properties
- Two-JAR deployment: hanlp-portable.jar and hanlp-solr-plugin.jar
- Drops into ${webapp}/WEB-INF/lib with a Solr restart
- Compatible with Solr StopFilterFactory for stop word removal
- Compatible with SynonymFilterFactory for index- and query-time synonyms
- Works alongside Solr LowerCaseFilterFactory in the analyzer chain
- Named entity recognition available through the underlying HanLP library
- Part-of-speech tagging available through the underlying HanLP library
- Handles ambiguous boundaries such as 商品和服务 without 和服 false positives
- Offline, self-hosted operation with no external API calls
- Apache Solr 5.x support, compatible with Lucene 5.x
About Hanlp Lucene Plugin
Hanlp Lucene Plugin brings HanLP's dictionary- and model-driven Chinese word segmentation into Apache Solr and Lucene. It's built for Java search engineers who need recall-friendly tokenization on Chinese text and don't want to hand-roll an analyzer. The plugin takes over from Solr's default text analyzers; you declare a fieldType in schema.xml, point index and query analyzers at com.hankcs.lucene.HanLPTokenizerFactory, and Solr indexes Chinese words instead of raw characters. Setup is deliberately small. Drop two JARs (hanlp-portable.jar, hanlp-solr-plugin.jar) into ${webapp}/WEB-INF/lib, optionally place hanlp.properties under the Solr resources folder for custom dictionaries, and restart. From there the fieldType is the whole interface: mark each field you want tokenized as type="text_cn" or reuse text_general so existing fields pick up HanLP automatically. The vendor's walkthrough uses Solr 5.2.1 and states 5.x compatibility. The design hinges on indexMode. Enable it on the index analyzer and HanLP performs full segmentation of long words, expanding terms like 中医药大学附属医院 into every sub-word it contains. Leave it off on the query analyzer, because the vendor explicitly warns that enabling it there breaks PhraseQuery. HanLP's disambiguation is the practical payoff: searching 和服 returns the price line, not the false hit a careless tokenizer makes by splitting 商品和服务 into 商品 和服 务. It slots into Solr's standard filter chain, so StopFilterFactory, SynonymFilterFactory and LowerCaseFilterFactory keep working alongside it. HanLP also carries named entity recognition and part-of-speech tagging. The scope is narrow by design: this is a self-hosted, offline Java plugin for Solr/Lucene, not a cloud tokenization API, and the closest alternatives in the same slot are IK Analyzer and Jieba Solr plugins.
Behind the Verdict
Reach for this when you're maintaining a Solr 5.x cluster with a Chinese corpus and the default analyzers are producing junk — either splitting real words or merging unrelated ones. The indexMode expansion is the genuinely useful part: it lets a query for 医院 or 中医药 pull documents whose text only contains 中医药大学附属医院, which is closer to how Chinese search users actually type than single-token matching. We'd also pick it over IK Analyzer or Jieba Solr when the disambiguation cases bite. The 商品和服务 versus 和服 example is the whole argument in miniature: a tokenizer that produces 商品 和服 务 returns a false positive on a 和服 query, and no amount of filter tuning fixes that after the fact. HanLP's dictionary plus model gets the boundary right before indexing. Where it bites: this is 2015 code. The tutorial anchors to Solr 5.2.1, docs are a blog post, and support is community-only via the GitHub repo. There's no managed option, no cloud API, and no adapter for non-Lucene stacks like Elasticsearch. If you're on Solr 8 or 9, you're testing an unmaintained plugin against a much newer server, and we wouldn't commit to that without a spike. Expect to know Java and Solr's schema. The failure modes are unglamorous but easy to hit — forgetting to set a field's type leaves it on the default analyzer and silently unsearchable, and turning on indexMode in the query analyzer damages PhraseQuery. The vendor flags both; read the tutorial before you touch schema.xml. The comparison to hold in your head: IK Analyzer and Jieba Solr plugins have broader version coverage and livelier communities, so if you're on a modern Solr they're the safer default. HanLP wins on tokenization quality for ambiguous Chinese and on the extra NLP machinery underneath. Pick based on which constraint you're actually
Researching Hanlp Lucene Plugin? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Hanlp Lucene Plugin actually fits — and what changes day-one when you adopt it.
You need to index Chinese product titles and descriptions with accurate segmentation to improve search recall and precision.
Outcome: You configure the HanLP tokenizer in schema.xml with index mode, set up custom dictionaries for product-specific terms, and see better search results with fewer false positives.
You need to tokenize Chinese legal documents to enable full-text search within a Solr 5.x deployment.
Outcome: You add the two JARs to your Solr webapp, configure a field type with HanLP, and leverage HanLP's NER and POS tagging to improve retrieval of case law and statutes.
You want to search Chinese titles and abstracts accurately in a Solr-based repository.
Outcome: You integrate the plugin, adjust query mode to preserve phrase searches, and use synonym filters to handle variant terminology, improving user satisfaction.
Use Cases
- Index Chinese product catalogs with accurate segmentation for e-commerce search
- Enable precise legal document retrieval by recognizing specialized terms
- Improve search relevance in academic journals by segmenting Chinese titles and abstracts
- Integrate with Solr to power multilingual search across Chinese and English content
Models Under the Hood
as of 2026-09-09
Limitations
- Lacks active commercial support and updates may lag behind HanLP's mainline.
- Requires manual Solr configuration.
- Does not provide an API service—only works within Lucene/Solr contexts.
as of 2026-08-31
Verification history
We have re-verified Hanlp Lucene Plugin 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Hanlp Lucene Plugin tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0
Ideal for
Search engineers and developers already using Solr 5.x or Lucene 5.x who need free, self-hosted Chinese tokenization without cost constraints.
What this tier adds
This is the only tier; it's free and open-source, with no paid support or premium features.
Where the pricing makes sense
The company stage and team size where Hanlp Lucene Plugin's pricing actually pencils out — and where peers do it cheaper.
This plugin is free, making it cost-attractive for small teams already invested in Solr 5.x. However, it lacks the ongoing development of paid alternatives like Elastic's official Chinese analyzer or commercial NLP APIs. For budget-conscious teams on legacy Solr, it's a good fit; for those needing modern support, consider paid options.
Setup time & first value
How long it actually takes to get something useful out of Hanlp Lucene Plugin — broken out by persona, not the marketing-page minute.
For a Solr 5.x expert, you can have HanLP tokenization working within 30 minutes: place two JARs, edit schema.xml, and restart Solr. For beginners, plan 1-2 hours due to Solr configuration and Java deployment nuances.
Switching to or from Hanlp Lucene Plugin
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From Solr's default analyzer: Replace tokenizer class in schema.xml with HanLPTokenizerFactory and enable index mode.
- ↗To IK Analyzer: Swap the tokenizer class to IKAnalyzerSolrFactory and adjust configuration files.
- ↗To Jieba Solr plugin: Replace tokenizer class with JiebaTokenizerFactory and update dictionary settings.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Hanlp Lucene Plugin”, and we withheld 6: 6 did not mention Hanlp Lucene Plugin. We are showing none, because we could not prove any of them are about Hanlp Lucene Plugin.
Official links
Featured Head-to-Head Comparisons
Hanlp Lucene Plugin vs Spider Cloud
Choose Hanlp Lucene Plugin if you're building a Chinese-language search engine on Solr/Lucene and need offline, customizable tokenization with NER. Choose Spider Cloud if you need fast, cost-effective web scraping and structured data extraction for AI agents and RAG pipelines, with a pay-as-you-go model and recent additions like Browser AI commands and data connectors. They solve completely different problems, so your choice depends on whether your focus is indexing Chinese text or fetching live web data.
Hanlp Lucene Plugin vs Voyage Ai
Hanlp Lucene Plugin is a free, specialized tool for Chinese tokenization in Solr/Lucene search engines, ideal for teams needing offline, customizable segmentation. Voyage AI targets enterprise RAG with domain-specific embeddings, long-context, and low-dimensional vectors but requires sales contact for pricing. Choose Hanlp if you build Chinese search with Solr; choose Voyage if you need high-accuracy retrieval for complex domains like finance or legal.
Hanlp Lucene Plugin vs Temporal Ai
Hanlp Lucene Plugin and Temporal AI solve completely different problems—one is a niche Chinese tokenizer for search indexing, the other a durable execution platform for building reliable workflows and AI agents. Choose Hanlp Lucene Plugin if you are maintaining a Solr/Lucene-based search engine requiring accurate Chinese segmentation with custom dictionaries. Choose Temporal AI if you need to orchestrate multi-step processes, AI agent pipelines, or microservices with automatic recovery, state persistence, and human-in-the-loop capabilities.
Popular in Developer Infrastructure
Temporal AI
Temporal is the durable execution platform for AI agents and long-running workflows that survive crashes, retries, and abandoned sessions.
Frequently Asked Questions
Categories
Used Hanlp Lucene Plugin? Help shape our editorial sentiment research.