Skip to content

perf: reduce Go FFI allocations, harden lifecycle concurrency, and optimize native FTS ingestion / 优化 Go FFI 分配、加固生命周期并发与原生全文索引写入 - #15

Open
sunhailin-Leo wants to merge 10 commits into
zvec-ai:mainfrom
sunhailin-Leo:codex/optimize-go-ffi-hotpaths

Conversation

@sunhailin-Leo

@sunhailin-Leo sunhailin-Leo commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

@chinaux

摘要 / Summary

  • 中文:三条主线:Go FFI 热路径降分配(cgo+purego,前序系列);原生 FTS 索引的插入提速与瘦身(经 zvec 子模块 pin 交付);并发正确性加固 + 单转换错误捕获 + 解码路径零分配(本轮新增两个 commit e8c273d/691c535,叠加在本 PR 头部)。
  • English: three threads: Go FFI hot-path allocation cuts (cgo+purego, earlier series), a native FTS ingestion optimization (delivered via the zvec submodule pin), and a new top-up series — concurrency-safety hardening, single-transition error capture, and a zero-alloc decode path (commits e8c273d/691c535).

改动 / Changes

  • 原生 FTS(zvec 子模块 → 43e62c6, fork 分支 perf/fts-store-positions-option, 基于 v0.7.0):FTS extra_params 新增索引级开关 {"store_positions": false}。关闭时 insert() 跳过每个 (term, doc) 的 $POS 位置写入——单文档 WriteBatch 中每 term 4 个 KV 写省 1 个;短语查询显式 InvalidArgument 拒绝而非静默错果;非布尔取值在 open() 校验失败。无新增 C 符号,老预编译库向前兼容(忽略该键)。另:insert() 复用 (term,doc_id) 键与 tf 缓冲。
  • Go FTS:NewFTSIndexParams 文档补充该开关;FTS 基准新增 NoPos 变体;新增两个集成测试。
  • Go FFI(前序系列):cgo/purego 查询结果打包进连续存储并释放结果指针数组;批量 Delete/Fetch 共用 C 主键缓冲;getter 减堆分配、新增可复用缓冲 GetVectorFP32FieldInto;purego 字符串数组合并;Darwin 原生堆回归测试。
  • 本轮新增 · 并发正确性与解码路径(e8c273d,cgo+purego 双后端):
    • 单转换错误捕获:zvec_go_status_t 返回结构把操作与线程局部 last-error 读取折进同一次原生调用(collection 的 insert/update/upsert/delete/delete_by_filter/query/multi_query/fetch/flush/close/destroy 与 doc 的字段访问全量迁移),消除 goroutine 在跨 cgo 调用迁移 OS 线程时的错误串扰窗口;热路径不再需要 lockErrorThread 逐调用锁线程(冷路径保留)。
    • 生命周期幂等:Collection.Close/Destroy 原子 CAS 门(并发双 Close 不再 double-free/SIGSEGV),所有操作经 liveness 门,闭合后统一返回新哨兵 ErrClosed;Doc.Destroy/FreeDocs 原子认领句柄,并发/重复销毁恰好释放一次。
    • 字段名 interning:sync.Map 缓存(上限 1024,溢出回退 fresh CString;PK/值绝不 intern),每字段访问省 1 次 native malloc+memcpy+free 与 2 次 cgo 转换。
    • GetPK 零拷贝(zvec_doc_get_pk_pointer);GetBinaryField/GetBinaryFieldInto 补上二进制字段读洞(此前只写不读);FreeDocs 经 zvec_go_docs_destroy shim 单次转换批量销毁。
    • 测试:-race 并发错误归属(8 goroutine × 20 轮,每条错误必须含自己的字段标记)、并发幂等 Close/Destroy、并发 FreeDocs、二进制往返、intern 上限回退。
  • 本轮新增 · 基准补全与守卫加固(691c535):BenchmarkCollectionQueryResultDecode(此前完全未测量的搜索解码半程:Query + 逐 doc GetPK/GetScore/Into + FreeDocs);BenchmarkCgoCallFloor_GetScore/_IsInitialized(单转换底价 ~29ns);首个 RunParallel 并发查询基准;Update/Upsert 批量写基准;doc 微基准去除 fmt.Sprintf 污染(SetPK/Add* 现在 0 allocs/op);Darwin 原生堆守卫改为稳态二次窗口测量 + 2MB 阈值(消除此前注记的 1MB 边界抖动)。

FTS 实测 / FTS measurements (native optimization)

条件 / Setup: 180,224 docs · varied 语料 · Batch512 · macOS arm64 · cgo source 模式,同机对照;本机高后台负载,时间为单次运行、含噪声;磁盘尺寸跨运行字节级稳定。

指标 / Metric v0.7.0 未修改 (n=1) 本 PR store_positions:false (n=3) Δ
Sum(Insert()) 167.4 s 129.4 s(中位;区间 113.9–148.5) -22.7%
Flush 后全目录 / dir after Flush 1,116.9 MB 706.4–719.0 MB -36.8% / -36.3% / -35.6%
Close 后全目录 / dir after Close 607.2 MB 302.5 MB -50.2%
Close 后 FTS rocksdb 目录 / FTS rocksdb dir 398.0 MB 93.3 MB -76.6%
  • 负对照 / negative control:同一 JSON 开关跑未修改 v0.7.0 库 → 167.4 s / 607.2 MB,与不开键位等价:native delta 全部来自原生改动。
  • 复现 / Reproduce: ZVEC_BENCH_FTS_ROWS=180224 ZVEC_BENCH_FTS_CORPUS=varied go test -tags integration -run '^$' -bench '^BenchmarkFTSIngestion/(FTS_Batch512|FTS_NoPos_Batch512)$' -benchtime=1x -count=1 -benchmem .

Go FFI 实测 / Go FFI measurements

Apple M3 Pro · Go 1.24.3 · zvec v0.7.0;三运行取中位;Go B/op 不含原生分配。

前序系列(前 → 后):

场景 / Case Before → After 分配 / Allocations
cgo Query, 128D/TopK100 127.6 → 126.6 µs/op 103 → 4 allocs/op
cgo FP32 Into, 128D 176.7 → 154.8 ns/op 2 → 0 allocs/op
cgo string getter 191.6 → 159.8 ns/op 3 → 1 allocs/op
purego Fetch, 1,000 unique keys 273,865 → 249,289 Go B/op 2,011 → 2,010 allocs/op

本轮新增(e8c273d 基线后实测,单转换底价 29.07 ns/op · 0 allocs):

场景 / Case 数值 说明
cgo DocSetPK 119.3 ns/op · 0 allocs 清 Sprintf 后纯 FFI
cgo DocAddStringField 217.6 ns/op · 0 allocs interning + 结构体返回消掉参数逃逸
cgo DocGetStringField 103.0 ns/op · 1 alloc 仅剩 GoStringN 结果字符串(地板)
解码往返 TopK100 235 µs/op · 105 allocs(TopK+5 地板) Query+逐 doc PK/score/Into+FreeDocs 全在计时内
解码往返 TopK10 55 µs/op · 15 allocs 同上
初步观察 Update 批量 81.4 vs Upsert 26.1 µs/key(≈3×) 新基准暴露,未深查(native 侧嫌疑大),后续跟进
  • 批量 Delete 的 C 字符串分配由每键一次降至每批一次。原生堆负对照(20,000 次 × 10 结果):故加稳态预热窗口后正常释放路径测量稳定通过(见下)。

验证 / Validation

  • zvec 原生 fts 列索引单测 74/74(含 5 个新用例)。
  • 本轮新增提交后复核:cgo source 全量套件 通过(含 FTS/iterator/multi-query/新并发测试);purego 全量套件通过(ZVEC_LIBRARY_PATH 加载重建库);新并发/二进制/intern 测试 -race 6/6 通过;go vet 双态、gofmt 通过。
  • CI:四平台 Source + 四平台 Vendor 8/8 绿(3d2077f);新提交 e8c273d/691c535 待 CI 复核。
  • 原生堆噪声注记已解决:TestCollectionQueryFetchNativeHeapDoesNotGrowPerResult 原单窗口测量把分配器一次性高水位增长与每结果泄漏混测(1.68 MB 负对照仍可检出真泄漏);691c535 增加稳态预热窗口,本开发机连续 4 次稳定通过。
  • ci.yml 与 .gitmodules 均未改动;pin 的 43e62c6 经 GitHub fork 网络可正常拉取。原生改动建议另提 PR 并入上游 zvec,届时把 gitlink 重指上游 commit。

后续 / Follow-ups

  • query.go/schema.go/zvec.go 冷路径仍走 toError+lockErrorThread(单 goroutine 构造型调用面),迁移 _ex 属机械扩展。
  • Update 批量比 Upsert 慢 ~3×:需双侧排查(Go 层同构,嫌疑在 native update path),必要时提上游 issue。
  • SearchQuery setter(SetFieldName/SetFilter/SetOutputFields)尚未接入 interning/单 arena 打包。

sunhailin-Leo and others added 7 commits September 24, 2026 14:29
…marks

Expose the zvec FTS index-level extra_params switch
{"store_positions": false} to Go users via the existing
NewFTSIndexParams extraParams argument (no new C symbols, so older
prebuilt libraries keep working).

- New integration tests: phrase queries keep working with the default
  (positions on); with the flag off, term/BM25 queries keep working,
  phrase queries are rejected, and the behavior survives Close/Open.
- FTS ingestion benchmark gains NoPos variants (Batch1/128/512/After)
  so insert-time and on-disk gains/losses are measurable per corpus.

Co-Authored-By: Claude <noreply@anthropic.com>
Bump the zvec submodule from v0.7.0 to 43e62c6 (fork branch
perf/fts-store-positions-option, based on v0.7.0) which adds the
native FTS extra_params switch {"store_positions": false}:
skip per-(term, doc) position writes to cut ingestion cost and index
size, and reject phrase queries with an explicit error.

.gitmodules stays untouched at alibaba/zvec: GitHub's fork network
shares object storage, so the pinned SHA is fetchable from the
submodule URL as-is. The native change is proposed upstream
separately; once merged, repoint the gitlink to the upstream commit.

Measured on 180,224 docs / varied corpus (macOS arm64, source mode):
sum of Insert() 167.4s -> 129.4s (median of 3 runs, -22.7% vs the
same benchmark on the unmodified library); closed collection total
607.2 MB -> 302.5 MB (-50.2%); closed FTS rocksdb dir 398.0 MB ->
93.3 MB (-76.6%). Sizes are byte-stable across runs. Positions-on
ingestion keeps identical on-disk output.

Co-Authored-By: Claude <noreply@anthropic.com>
@sunhailin-Leo
sunhailin-Leo force-pushed the codex/optimize-go-ffi-hotpaths branch from 13bc090 to 3d2077f Compare September 25, 2026 07:02
@sunhailin-Leo
sunhailin-Leo marked this pull request as ready for review September 25, 2026 08:43
@sunhailin-Leo
sunhailin-Leo marked this pull request as draft September 25, 2026 08:44
@sunhailin-Leo
sunhailin-Leo marked this pull request as ready for review September 25, 2026 08:44
sunhailin-Leo and others added 2 commits September 28, 2026 15:16
…stroy, zero-alloc field accessors

cgo backend (doc.go / collection.go / errors.go):
- Fold each operation and its thread-local last-error fetch into ONE
  native call via zvec_go_status_t-returning wrappers
  (zvec_go_collection_{insert,update,upsert,delete,delete_by_filter,
  query,multi_query,fetch,flush,close,destroy}_ex and
  zvec_go_doc_{add_field_by_value,set_field_null,remove_field,
  get_field_names}_ex / zvec_go_get_field_ex). This removes the error
  cross-wiring a goroutine migrating OS threads between FFI calls could
  observe under concurrent failure load, and drops the per-call
  LockOSThread pin now needed only by unmigrated cold-path sites.
- Field names are interned (sync.Map, cap 1024, fresh-CString fallback):
  per-field accessors stop paying a native malloc+memcpy+free pair and
  two cgo transitions per call. PKs/values are never interned.
- GetPK switches to the zero-copy zvec_doc_get_pk_pointer; GetFieldNames
  returns via struct; basic getters read the value pointer directly.
- FreeDocs batch-destroys handles via a zvec_go_docs_destroy C shim
  (never frees the Go-backed array itself); Doc.Destroy claims its
  handle atomically.
- Collection Close/Destroy gain an atomic.Bool CAS gate (idempotent,
  double-free-free); all operations check liveness and return the new
  ErrClosed sentinel.
- GetBinaryField / GetBinaryFieldInto restore binary round-trip (binary
  fields were previously write-only).

purego backend (purego_api.go): same lifecycle semantics — CAS-gated
Close/Destroy, live() gates + ErrClosed on every operation, atomic Doc
handle claim, GetBinaryField/GetBinaryFieldInto parity.

Tests: -race concurrent error attribution (8 goroutines x 20 rounds,
each error must name its own field marker), concurrent idempotent
Close/Destroy, concurrent FreeDocs/Destroy, binary round-trip, intern
cap fallback.

Measured (M3 Pro, cgo): doc microbenchmarks after fmt.Sprintf cleanup —
AddStringField 217ns 0 allocs, GetStringField 103ns 1 alloc, SetPK
119ns 0 allocs; full decode round trip TopK100 = 105 allocs/op (the
TopK+5 floor), cgo transition floor ~29ns.

Co-Authored-By: Claude <noreply@anthropic.com>
…n native-heap guard

- BenchmarkCollectionQueryResultDecode: the previously unmeasured
  user-visible half of a search (Query + per-doc GetPK/GetScore/
  GetVectorFP32FieldInto into a shared buffer + FreeDocs), TopK 10/100,
  ns/doc metric. QueryPhases deliberately excludes decode from its
  timed regions.
- BenchmarkCgoCallFloor_GetScore/_IsInitialized: per-transition floor.
- BenchmarkCollectionQueryParallel: first RunParallel benchmark in the
  package; BenchmarkCollectionUpdateBatch/UpsertBatch: batched write
  paths over a fixed corpus (exposes Update ~3x slower than Upsert).
- Clean fmt.Sprintf out of the doc microbenchmark loops (SetPK/Add*
  now measure pure FFI: 0 allocs/op).
- Harden the Darwin native-heap guard: prime the allocator with a
  20,000-call warmup window before measuring (single-window drift
  conflates one-time malloc high-water growth with per-result leaks and
  flakes with system memory pressure) and document the 2MB steady-state
  threshold.

Co-Authored-By: Claude <noreply@anthropic.com>
@sunhailin-Leo sunhailin-Leo changed the title perf: reduce Go FFI allocations and optimize native FTS ingestion / 优化 Go FFI 分配与原生全文索引写入 perf: reduce Go FFI allocations, harden lifecycle concurrency, and optimize native FTS ingestion / 优化 Go FFI 分配、加固生命周期并发与原生全文索引写入 Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant