Install
npm install @pacote/bloom-search
The package ships as ES modules and CommonJS, with TypeScript types. It has no stemmer of its own; the examples below use the stemmer package.
Describe the index
Create the instance in one module, so the build script and the browser cannot drift apart.
// search.ts
import { BloomSearch } from '@pacote/bloom-search'
import { stemmer } from 'stemmer'
type Post = { url: string; title: string; body: string }
export const search = new BloomSearch<Post, 'url' | 'title', 'title' | 'body'>({
fields: { title: 3, body: 1 }, // a title word counts three times
summary: ['url', 'title'], // everything the index will return
stemmer,
})
The three type parameters are the document, the fields kept as summary, and the fields indexed.
Build it
At build time, index your documents and write the index out.
import { writeFileSync } from 'node:fs'
import { search } from './search'
for (const post of posts) search.add(post.url, post)
writeFileSync('public/search-index.json', JSON.stringify(search.index))
JSON works. MessagePack gives a smaller file; the demos on this site use it.
Search in the browser
Load the index once, then reuse the instance.
import { search } from './search'
search.load(await (await fetch('/search-index.json')).json())
search.search('whale +ahab -pequod')
// => [{ url: '/moby-dick', title: 'Moby-Dick' }, ...]
Build and browser must agree on the stemmer, the seed and the tokenizer. Read the limits before you choose it.
Queries
| Write | Meaning |
|---|---|
whale |
Optional. Matching documents rank higher. |
+whale |
Required. |
-whale |
Excluded. |
"white whale" |
Exact phrase. Needs ngrams of 2 or more,
up to the phrase length.
|
Terms are lowercased and run through your stemmer, exactly like document words.
Options
| Option | Default | What it does |
|---|---|---|
fields |
required | Fields to index: an array, or an object of field names and weights. |
summary |
required | Fields kept in the index and returned in results. Keep it to what identifies a document. |
errorRate |
0.0001 | Target false-positive rate per filter. Lower is more reliable and larger. |
minSize |
0 | Minimum term count used to size a filter. Raise it for tiny documents. |
ngrams |
1 | Also index word sequences up to this length, which phrase search needs. Roughly multiplies index size by n. |
termFrequencyBuckets |
[1, 2, 3, 4, 8, 16, 32, 64] | Frequency thresholds that split a document into filters. |
seed |
0x00c0ffee | Hash seed. Build and browser must agree. |
preprocess |
String | Turns a field value into text before tokenizing, for example to strip HTML. |
tokenizer |
lowercase, split on whitespace and hyphens | Splits text into tokens. |
stemmer |
none | Maps a token to its stem. Applied to documents and to queries. |
stopwords |
index everything | Predicate: return true to index a token, false to drop it. |
Methods
| Member | What it does |
|---|---|
add(ref, document, language?) |
Indexes a document under a unique reference. Adding the same reference again replaces it. |
remove(ref) |
Drops a document from the index. |
search(query, language?) |
Returns the summaries of matching documents, best first. |
index |
The serializable index: a version number and every document’s summary and filters. |
load(index) |
Replaces the index with one built elsewhere. Throws if the schema version differs. |
toJSON() |
Options and index together. |
Every option and method is documented in the API reference.