Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .jules/bolt.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
## 2023-11-20 - Memoizing regex compilations in ConText
**Learning:** In the openmed clinical text processing pipeline, deterministic regex compilations (via `_compiled_context_lexicon` in `openmed.clinical.context`) create significant performance bottlenecks when repeatedly evaluated across large document streams. Compiling all ConText regexes every time context cues are evaluated is extremely slow because regex parsing adds immense overhead per execution for static definitions.
**Action:** When working on NLP pipelines, always memoize deterministic regex and lexicon compilations, such as using `@functools.lru_cache` to cache instance instantiations that depend merely on a simple string (e.g. language). Ensure that return values are immutable so callers don't accidentally mutate cached instances.
2 changes: 2 additions & 0 deletions openmed/openmed/clinical/context.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@

from __future__ import annotations

import functools
import re
from collections.abc import Iterable, Iterator, Mapping, Sequence
from dataclasses import dataclass, replace
Expand Down Expand Up @@ -154,6 +155,7 @@ class _CompiledContextLexicon:
backward_context_cues: frozenset[str]


@functools.lru_cache(maxsize=16)
def _compiled_context_lexicon(language: str | None = None) -> _CompiledContextLexicon:
lexicon = get_clinical_cue_lexicon(language)
token_boundaries = lexicon.token_boundaries
Expand Down