Releases: openclaw/discrawl
Releases · openclaw/discrawl
Release list
v0.4.1
Fixes
- existing archives that already report schema version 2 now self-heal missing embedding tables and columns before 0.4.x sync/update commands continue.
v0.4.0
Changes
- semantic message search now ranks across the full compatible local vector set instead of only the newest candidate window. (#36) Thanks @GaosCode.
- hybrid message search now fuses FTS with local semantic vectors while avoiding embedding-provider calls when no local vectors exist. (#37) Thanks @GaosCode.
- local embedding providers now support OpenAI-compatible endpoints, Ollama, and llama.cpp, and
doctorcan probe the configured provider before you queue vectors embednow drains the queued embedding backlog in bounded batches, requeues safely on provider throttling, and drops stale stored vectors when messages no longer have embeddable content- Git snapshot publishing can now opt in to backing up generated embedding vectors with
--with-embeddingswhile still keeping embedding queue state local. - Git-backed snapshot imports are now much faster on large archives by using import-only SQLite pragmas and bulk-load FTS5 settings during search index rebuilds
messagesandmentionsnow use composite read-path indexes so larger archives spend less time sorting/filtering common guild, channel, and author queries
Fixes
- normalized message text is now sanitized before it reaches SQLite and FTS5, repairing malformed UTF-8 and stripping invisible/control-character noise that can poison search content
- Git-backed snapshots now keep embedding queue state and generated vectors local to each archive, so subscribers no longer inherit misleading embedding backlog metadata. (#38) Thanks @GaosCode.
Docs
- docs now cover semantic and hybrid search setup, embedding privacy, Git snapshot behavior, and local vector rebuilds. (#39) Thanks @GaosCode.
Tests
- Git embedding snapshot export/import now has CLI, share-package, and Docker E2E coverage.
- total Go test coverage now reaches the 85% line.
v0.3.0
sync --allnow bypassesdefault_guild_idso one run can fan out across every discovered guild without clearing the single-guild default firstsync --fullno longer aborts when forum thread discovery hits Discord403 Missing Access; inaccessible channels are skipped and marked unavailable while accessible channels continue syncing- startup now validates and stamps SQLite schema version via
PRAGMA user_version, and fails fast if the local DB schema is newer than the running binary - git-backed archive sharing can now export/import compressed JSONL snapshots with manifests, subscribe to a Git repo as the data source, and run in git-only mode without Discord credentials
messages,search, and reports can automatically refresh stale git-backed data, preferring the Git snapshot before falling back to live Discord when both sources are configured- the Discord backup publisher workflow now syncs latest messages, publishes the archive to a private GitHub repo, serializes concurrent runs, validates required secrets, and skips the member crawl for faster updates
- the backup report workflow now updates README activity stats from the backup action and keeps those queries bounded with process timeouts
sync --latest-onlyadds a lightweight refresh path for checking recent Discord messages without doing a full historical crawl- repository imports now skip expensive rebuilds when the snapshot manifest is already current, and GitHub Actions persist the warmed SQLite database across runs
- the Docker git-source smoke test now verifies that a fresh install can subscribe to a repository-only archive and query messages, SQL, and reports
- CI now uses Go 1.26.2,
actions/setup-go6.4.0, cache actions 5.0.5, Node 24 for report generation, and refreshed SQLite dependencies
v0.2.0
- much faster
sync --fullbehavior on large archives: incomplete backfills are auto-batched, active-thread discovery is more precise, and steady-state refreshes avoid re-scanning every archived thread once history is already complete sync --sincenow reliably honors the cutoff during bootstrap and full-history backfill, while still allowing a latersync --fullwithout--sinceto continue older history- full-sync progress is more resilient: slow member crawls no longer hold message sync hostage, and stale unavailable-channel markers are cleared so recovered channels can sync again
- offline member-profile search is now much richer:
members searchmatches archived profile fields in addition to names members shownow accepts either Discord IDs or queries and can include recent messages plus message stats for the resolved member- archived profile extraction now surfaces stored fields like
bio,pronouns,location,website,x,github, and discovered URLs when present messages --synccan do a blocking pre-query refresh for the matching channel or guild scope before reading the local archivemessages --hoursadds recent-hour slices without manual RFC3339 timestampsmessages --lastreturns the newest matching rows while still printing them oldest-to-newest
v0.1.0
- initial public release of
discrawl - multi-guild Discord crawler with single-guild default UX
- local SQLite archive with FTS5 search
- commands:
init,sync,tail,search,messages,mentions,sql,members,channels,status,doctor - env-based bot token discovery
- resumable full-history sync, live gateway tailing, repair sync loop, targeted channel sync
- attachment-text indexing for small text-like uploads
- structured user and role mention indexing/querying
- empty-message filtering based on real searchable/displayable content instead of raw body only
- CI with lint, tests, secret scanning, and coverage enforcement
- release plumbing via GoReleaser, GitHub Actions, and Homebrew tap packaging
- sync correctness fixes for empty channels, inaccessible channels, unknown channels, and large-channel resume behavior
- SQLite/FTS performance fixes for backfill throughput and lower write amplification