
类型:数据库
简介:实时且性能出色的向量数据库,专门针对大规模向量搜索进行优化。
Pinecone 稀疏向量词法检索,专门针对上游生成的令牌权重稀疏向量做精准关键词召回,适配 pinecone-sparse-english-v0 等各类自定义稀疏编码模型,由业务侧直接管控稀疏向量生成格式。
一、词法检索适用场景对比
通用商品搜索、内容推荐、文档问答等常规文本检索,优先使用BM25全文搜索,基于开启全文检索的字符串字段,依托Lucene语法与BM25相关性打分排序。同一份多字段索引,后续也可追加稠密向量、稀疏向量排序能力。
稀疏向量词法检索属于独立检索模式,适合已有自研稀疏编码,提前生成好词袋权重向量的业务场景。稀疏向量维度极高,全文仅有少量非零数值,维度对应词典词汇,数值代表词语在文档中的重要程度,词语独立打分后累加排序,相似度越高排名越靠前。
二、文本方式稀疏向量检索
文本直接检索仅支持开启集成嵌入能力的稀疏向量索引。传入查询语句后,Pinecone自动调用索引绑定模型,将文本转为稀疏向量并完成相似度检索。
接口核心参数:
- namespace:查询命名空间,默认空间填写default
- query.inputs.text:用户查询关键词文本
- query.top_k:返回匹配结果条数
- match_terms:可选,强制命中指定关键词列表
- fields:可选,自定义返回结果字段
Java 调用示例代码:
import io.pinecone.clients.Index;import io.pinecone.configs.PineconeConfig;import io.pinecone.configs.PineconeConnection;import org.openapitools.db_data.client.ApiException;import org.openapitools.db_data.client.model.SearchRecordsResponse;
import java.util.*;
public class SearchText { public static void main(String[] args) throws ApiException { PineconeConfig config = new PineconeConfig("YOUR_API_KEY"); config.setHost("INDEX_HOST"); PineconeConnection connection = new PineconeConnection(config);
Index index = new Index(config, connection, "integrated-sparse-java");
String query = "What is AAPL's outlook, considering both product launches and market conditions?"; List fields = new ArrayList<>(); fields.add("category"); fields.add("chunk_text");
SearchRecordsResponse recordsResponse = index.searchRecordsByText(query, "example-namespace", fields, 3, null, null); System.out.println(recordsResponse); }}
返回结果按照相似度分数从高到低排序,同时消耗对应读取单元与嵌入 Token 数量。
三、原生稀疏向量检索
直接传入自定义稀疏向量索引与权重数值,执行精准词法匹配检索,适配所有自研稀疏编码场景。
接口核心参数:
- namespace:指定查询命名空间
- sparse_vector:稀疏向量索引下标+对应权重数值
- top_k:返回结果数量
- include_values:是否返回匹配向量原值,默认关闭
- include_metadata:是否返回文档元数据,默认关闭
当 top_k 大于 1000 时,建议关闭向量与元数据返回,仅保留 ID 与分数,大幅降低查询延迟、提升检索性能。
Java 稀疏向量查询示例:
import io.pinecone.clients.Pinecone;import io.pinecone.unsigned_indices_model.QueryResponseWithUnsignedIndices;import io.pinecone.clients.Index;
import java.util.*;
public class SearchSparseIndex { public static void main(String[] args) throws InterruptedException { Pinecone pinecone = new Pinecone.Builder("YOUR_API_KEY").build();
String indexName = "docs-example"; Index index = pinecone.getIndexConnection(indexName);
List sparseIndices = Arrays.asList( 767227209L, 1640781426L, 1690623792L, 2021799277L, 2152645940L, 2295025838L, 2443437770L, 2779594451L, 2956155693L, 3476647774L, 3818127854L, 428309169L); List sparseValues = Arrays.asList( 1.0f, 1.0f, 1.0f, 1.0f, 1.0f, 1.0f, 1.0f, 1.0f, 、1.0f, 1.0f, 1.0f, 1.0f);
QueryResponseWithUnsignedIndices queryResponse = index.query(3, null, sparseIndices, sparseValues, null, "example-namespace", null, false, true); System.out.println(queryResponse); }}
四、根据文档ID相似检索
传入已有记录ID,Pinecone自动调取该ID对应的稀疏向量作为查询条件,召回库内高度相似文档,适合相似文档推荐、关联内容匹配场景。
参数配置与稀疏向量检索一致,大批量高 top_k 查询同样建议关闭向量、元数据返回,优化接口响应速度。
JavaID 检索示例:
import io.pinecone.clients.Index;import io.pinecone.configs.PineconeConfig;import io.pinecone.configs.PineconeConnection;import io.pinecone.unsigned_indices_model.QueryResponseWithUnsignedIndices;
public class QueryExample { public static void main(String[] args) { PineconeConfig config = new PineconeConfig("YOUR_API_KEY"); config.setHost("INDEX_HOST"); PineconeConnection connection = new PineconeConnection(config); Index index = new Index(connection, "INDEX_NAME"); QueryResponseWithUnsignedIndices queryRespone = index.queryByVectorId(3, "rec2", "example-namespace", null, false, true); System.out.println(queryResponse); }}
五、强制关键词命中过滤
该功能处于公开预览,仅支持2025-10版本API接口。文本检索时可设置强制必含关键词,所有结果必须全部包含指定词汇,大幅提升检索精准度。
适用场景:实体精准筛选、领域术语限定、内容合规校验、股票行业定向检索、关键概念锁定。
使用方式:添加 match_terms 参数,传入必选词语 + 匹配策略,目前仅支持 all 策略,也就是所有关键词必须同时存在。
Curl 请求示例:
PINECONE_API_KEY="YOUR_API_KEY"INDEX_HOST="INDEX_HOST"
curl "https://$INDEX_HOST/records/namespaces/example-namespace/search" -H "Content-Type: application/json" -H "Api-Key: $PINECONE_API_KEY" -H "X-Pinecone-Api-Version: unstable" -d '{ "query": { "inputs": { "text": "What is the current outlook for Tesla stock performance?" }, "top_k": 3, "match_terms": { "terms": ["Tesla", "stock"], "strategy": "all" } }, "fields": ["chunk_text"] }'
未开启关键词过滤时,可能出现只含特斯拉、只含股票、两者都不相关的无效结果,过滤后仅返回同时包含双关键词的精准内容。
六、稀疏词法检索使用限制
- 关键词强制过滤仅支持集成嵌入类型稀疏索引
- 关键词筛选属于查询后二次处理,top_k以外匹配内容不会被过滤展示
- 不支持精确短语顺序匹配,仅匹配独立词汇
- 检索自动大小写归一化,不区分英文大小写
- 向量存储于对象存储,大量返回向量数据会明显增加接口延迟

