From 3238fba5e67f24b44ef03facb0e4f502bb05c6dd Mon Sep 17 00:00:00 2001 From: Teingi Date: Sat, 5 Sep 2026 01:55:54 +0800 Subject: [PATCH 1/3] docs(rfc): define unified artifact tags --- docs/en/rfcs/0000_artifact_tags.md | 687 +++++++++++++++++++++++++++++ docs/zh/rfcs/0000_artifact_tags.md | 646 +++++++++++++++++++++++++++ 2 files changed, 1333 insertions(+) create mode 100644 docs/en/rfcs/0000_artifact_tags.md create mode 100644 docs/zh/rfcs/0000_artifact_tags.md diff --git a/docs/en/rfcs/0000_artifact_tags.md b/docs/en/rfcs/0000_artifact_tags.md new file mode 100644 index 000000000..bc952ef38 --- /dev/null +++ b/docs/en/rfcs/0000_artifact_tags.md @@ -0,0 +1,687 @@ +- Proposal Name: `unified_artifact_tags` +- Start Date: 2026-09-05 +- RFC PR: [oceanbase/powercontext#0000](https://github.com/oceanbase/powercontext/pull/0000) +- Related Discussion: [oceanbase/powercontext#1466](https://github.com/oceanbase/powercontext/issues/1466) +- Related RFCs: [RFC 0014](0014_memory_layer_design.md), [RFC 0019](0019_local_source_memory_runtime.md), + [RFC 0048](0048_handoff_artifact.md), [RFC 0051](0051_experience_skill_artifact_families.md), + [RFC 1396](1396_handoff_access_control.md), and [RFC 1437](1437_source_artifact_rest_api.md) + +# Summary + +This RFC adds scope-local custom tags to PowerContext-managed content. A user can attach tags such as `customer-a`, +`cockpit`, or `verified` to a managed Experience, Skill, Handoff, whole Memory Artifact, or individual Memory entry, +then retrieve current resources by exact tag membership. + +The feature uses one `pc_artifact_tags` assignment table for every supported target. Each row retains the owning +Artifact identity and identifies either that Artifact or one logical Memory entry inside it. This preserves a uniform +product and API model without pretending that a Memory entry is an independent Artifact. + +Tags are mutable catalog attributes of a logical target. They are not Artifact content, Source metadata, lineage, +authorization policy, lifecycle state, or prompt instructions. Tag changes do not create Artifact Revisions, alter +content digests, rebuild embeddings, or change exact citations. Tags remain local to their owner Scope and are not +copied by Artifact publication unless a later contract explicitly requests that behavior. + +The first delivery provides tag read and compare-and-swap replacement, exact `all` and `any` tag queries, optional tag +filters on current-resource listing and Memory search, and a minimal Dashboard editor and filter. It deliberately does +not add tag hierarchies, colors, aliases, automatic tagging, tag-based authorization, or historical tag snapshots. + +# Motivation + +Teams accumulate many Memory entries, Experiences, and Skills in the same Scope. Text and semantic search answer +"which content resembles this query?" but do not reliably answer organizational questions such as: + +- Which entries belong to customer A? +- Which Skills have been validated for the automotive cockpit environment? +- Which Experiences describe release operations rather than inference behavior? +- Which retired or inactive resources still belong to a compliance review set? + +Users can encode some of this information in content, but doing so mixes classification with the knowledge itself. +Changing a classification would then create a new immutable Revision, change the content digest, and potentially +rebuild a search projection even though the reusable content did not change. + +Existing fields named `metadata` do not provide a shared solution. `ContentSource.metadata` contains provenance and +behavior-bearing fields such as `kind`. Managed Skill metadata belongs to exact Skill content and participates in its +search projection. Treating either field as a generic mutable tag bag would blur ownership, version, indexing, and +security semantics. + +The persistence model also has two observable target granularities: + +- Experience, Skill, Handoff, and a whole Memory are logical Artifact lifecycles selected by `artifact_id`; +- the Memory shown to a user as one remembered fact or preference is a logical entry selected by `entry_id` inside the + Scope's Memory Artifact. + +Putting one JSON tag column only on `pc_artifact_heads` would give every entry in the standard one-Memory-per-Scope +profile the same tag set. Adding separate family-specific tag stores would preserve granularity but fragment the API +and cross-family query path. PowerContext needs one explicit assignment model that handles both target shapes. + +# Guide-level explanation + +## User model + +A tag is a user-authored string attached to one current logical target in one Scope. It is useful for exact grouping +and filtering. It does not make a claim true, approve an Artifact, grant access, or instruct an Agent. + +The same visible behavior applies to every supported resource: + +```text +Scope: vehicle-assistant + +Memory entry: "The driver prefers 24 C in winter" +Tags: [customer-a, cockpit, preference] + +Experience: "Regenerate the Client after editing OpenAPI" +Tags: [release, verified] + +Skill: "vehicle-log-triage" +Tags: [customer-a, diagnostics] + +Handoff: "Complete the cockpit latency investigation" +Tags: [customer-a, in-progress] +``` + +Users may also tag the whole Memory Artifact when the classification applies to the collection rather than one entry. +The UI must distinguish `Memory` from `Memory entry` so that a collection tag is not mistaken for an entry tag. + +## Add and remove tags + +The Dashboard displays tags as inert, escaped chips. A user with write authority opens a resource, edits the complete +tag set, and saves it. Saving is conditional on the tag set observed by the editor so that two editors cannot silently +overwrite each other. + +For example, reading the current set returns an opaque `ETag` response header and this body: + +```json +{ + "tags": ["cockpit", "customer-a"], + "tag_digest": "sha256:92f..." +} +``` + +The user replaces it with: + +```json +{ + "tags": ["cockpit", "customer-a", "verified"] +} +``` + +The client sends the observed ETag in `If-Match`. It does not derive a precondition from `tag_digest`. + +Replacing a set with the same normalized tags is idempotent. Replacing it with an empty array removes every tag but +does not delete, retire, or revise the target. + +## Retrieve by tags + +Tag matching is exact after normalization. A filter for `customer-a` does not match `customer-a-archive`, and a tag +named `release/security` has no implicit parent named `release`. + +Multiple tags have an explicit match mode: + +- `all` selects targets that have every requested tag; +- `any` selects targets that have at least one requested tag. + +For example: + +```json +{ + "families": ["memory", "experience", "skill"], + "target_types": ["artifact", "memory_entry"], + "tags": ["customer-a", "verified"], + "match": "all", + "limit": 50 +} +``` + +The response returns logical targets together with current exact references. An Artifact result includes its current +`ArtifactReference`; a Memory entry result includes its current `MemoryCitation`. The exact references let a caller +read content without resolving `latest` a second time. + +Text retrieval can combine a query with a tag filter. PowerContext applies the tag filter to the eligible candidate +set before FTS, vector top-k selection, fusion, or reranking. Filtering an already truncated top-k result is incorrect +because a relevant tagged item may have been excluded before the filter ran. + +## Revision behavior + +Tags follow logical identity: + +```text +Experience exp-1 Revision 1 --\ +Experience exp-1 Revision 2 ----> tags for logical exp-1 + +Memory entry entry-1 Version 1 --\ +Memory entry entry-1 Version 2 ----> tags for logical entry-1 +``` + +Revising `exp-1` or `entry-1` therefore preserves its tags. Tagging either target does not create a content Revision. +An exact historical Artifact or Memory citation remains a statement about immutable content and evidence; it does not +implicitly acquire a historical tag snapshot. + +## Scope and publication behavior + +Tags belong to the Scope in which they were assigned. Publishing or copying an Artifact to another Scope creates the +destination logical Artifact with an empty tag set. This default avoids leaking customer names, internal workflow +categories, or other local classifications. A caller may assign destination tags explicitly after publication. + +# Reference-level explanation + +## Goals + +The first implementation must: + +- provide one tag model for managed Artifacts and logical Memory entries; +- persist all assignments in one table; +- keep tag mutation independent from immutable Artifact and Memory-entry content; +- support exact `all` and `any` filtering with deterministic normalization; +- combine tags correctly with current Artifact listing and Memory retrieval; +- preserve Scope isolation, target visibility, and existing lifecycle filters; +- provide concurrency-safe, idempotent complete-set replacement; and +- expose enough current exact identity for a query result to be resolved safely. + +## Non-goals + +The first implementation does not define: + +- tag colors, descriptions, aliases, hierarchy, inheritance, or a tag-definition catalog; +- automatic tag generation by a model, Source metadata, or content keywords; +- tags as authorization, approval, trust, lifecycle, routing, or retention policy; +- historical tag snapshots or tag assignment audit history; +- tag propagation through lineage, publication, fork, import, or Handoff evidence; +- revision-specific tags; +- MCP tools that let an Agent mutate tags automatically; or +- arbitrary key/value Artifact metadata. + +## Terminology and target identity + +`ArtifactTagTarget` is a discriminated union: + +```text +ArtifactTagTarget = + ArtifactTarget { + type: "artifact", + family: string, + artifact_id: string + } + | MemoryEntryTarget { + type: "memory_entry", + family: "memory", + artifact_id: string, + entry_id: string + } +``` + +An `ArtifactTarget` identifies a logical Artifact lifecycle and intentionally omits `revision`. A +`MemoryEntryTarget` identifies one logical Memory entry and intentionally omits both Memory Revision and +`entry_version_id`. The containing `artifact_id` remains explicit because `entry_id` is scoped by its Memory Artifact. + +The canonical persistence form uses `target_id`: + +| Target | `family` | `artifact_id` | `target_type` | `target_id` | +| --- | --- | --- | --- | --- | +| Experience | `experience` | Experience ID | `artifact` | same Experience ID | +| Skill | `skill` | Skill ID | `artifact` | same Skill ID | +| Handoff | `handoff` | Handoff ID | `artifact` | same Handoff ID | +| Whole Memory | `memory` | Memory ID | `artifact` | same Memory ID | +| Memory entry | `memory` | containing Memory ID | `memory_entry` | entry ID | + +New nested resource kinds must not reuse `memory_entry` or overload `target_id`. They require an explicit target type +and validation rule in a follow-up contract. + +## Tag value and normalization + +A submitted tag must satisfy all of the following: + +- it is a Unicode string between 1 and 64 Unicode code points; +- it has no leading or trailing whitespace; +- it contains no Unicode control, surrogate, or unassigned code point; +- its normalized key is at most 128 Unicode code points; and +- the complete submitted set contains at most 32 tags. + +PowerContext preserves the submitted value as `tag` for display. It computes `tag_key` by applying Unicode NFC and +then Unicode default case folding. It does not collapse internal whitespace, split punctuation, parse `/`, translate, +stem, or infer hierarchy. Both `tag` and `tag_key` are validated after normalization. + +Two submitted tags with the same `tag_key` are duplicates and the complete request is rejected. The Server does not +silently choose a display spelling. A later successful replacement may change only the preserved display spelling +while retaining the same `tag_key`. + +Tag sorting is ascending by UTF-8 bytes of `tag_key`. This order is used in responses and digest calculation so that +database collation does not change public behavior. + +## Persistence + +The shared relational schema adds exactly one business table: + +```text +pc_artifact_tags + scope_id identity string, not null + family identity string, not null + artifact_id identity string, not null + target_type identity string, not null + target_id identity string, not null + tag_key identity string, not null + tag display string, not null + assigned_at UTC timestamp, not null + + primary key ( + scope_id, + family, + artifact_id, + target_type, + target_id, + tag_key + ) + + foreign key (scope_id, family, artifact_id) + references pc_artifact_heads (scope_id, family, artifact_id) + on delete cascade + + check target_type in ('artifact', 'memory_entry') + check target_type != 'artifact' or target_id = artifact_id + check target_type != 'memory_entry' or family = 'memory' +``` + +The table has these secondary indexes: + +```text +(scope_id, family, tag_key, target_type, artifact_id, target_id) +(scope_id, tag_key, family, target_type, artifact_id, target_id) +``` + +The primary key supports loading one target's tags. The first secondary index supports family-specific filtering; the +second supports a cross-family Scope query. Implementations must use binary identity comparison or application-built +`tag_key` values rather than depending on database-default case or locale collation. + +One table cannot express conditional foreign keys to both `pc_artifact_heads` and `pc_memory_entry_heads`. The owning +Artifact foreign key is enforced for every row. For `memory_entry`, the repository must additionally lock and validate +the current `(scope_id, memory_artifact_id, entry_id)` head in the same transaction before changing assignments. The +transaction rejects a missing entry. This is an explicit application invariant, not a best-effort cleanup rule. + +Inactive Memory entries and deprecated or retired Artifacts retain their assignments. Tag queries apply the existing +visibility and lifecycle selection before returning results. A caller with the target's write authority may reorganize +tags without reactivating or revising content. + +## Tag set and digest + +`ArtifactTagSet` contains: + +```json +{ + "scope_id": "vehicle-assistant", + "target": { + "type": "memory_entry", + "family": "memory", + "artifact_id": "memory", + "entry_id": "mem_ent_123" + }, + "tags": ["cockpit", "customer-a"], + "tag_digest": "sha256:..." +} +``` + +`tag_digest` is the SHA-256 digest of RFC 8785 canonical JSON for the object `{"tags": [...]}`, where tags are in +canonical `tag_key` order and each array item is the preserved display string. The empty set has a stable digest. It +is a content checksum for the tag set, not a client-visible compare-and-swap token, Artifact content digest, Memory +entry content hash, or authorization generation. + +HTTP operations use an opaque ETag that binds the complete logical target identity and `tag_digest`. Clients must not +assume that the ETag equals, embeds, or can be reconstructed from `tag_digest`. + +## Repository contract and transactions + +The tag repository exposes three operations: + +```text +get(scope_id, target) -> ArtifactTagSet + +replace( + scope_id, + target, + expected_tag_digest, + tags +) -> ArtifactTagSet + +query( + scope_id, + tags, + match, + families, + target_types, + lifecycle_selection, + limit, + cursor +) -> ArtifactTagPage +``` + +`replace` performs the following steps in one transaction: + +1. Resolve and authorize the target without returning hidden existence details. +2. Lock the owning Artifact head. For a Memory entry, also lock and validate its current entry head. +3. Load the complete current tag set and calculate its digest. +4. Reject a mismatched expected digest as a failed precondition. +5. Validate and normalize the complete replacement set. +6. Delete assignments absent from the replacement and insert or update the remaining display values. +7. Return the canonical set and new digest. + +Locking the existing owning head serializes replacement even when the current tag set is empty. Implementations must +not rely on locking zero assignment rows or on a process-local lock, because neither provides the required distributed +compare-and-swap behavior. + +If the current canonical set equals the requested set, `replace` is an idempotent success and changes no rows or +`assigned_at` values. + +## HTTP contract + +The first contract adds five operations below the Scope resource tree established by RFC 1437: + +| Method | Path | operationId | Purpose | +| --- | --- | --- | --- | +| `GET` | `/v1/scopes/{scope_id}/artifacts/{family}/{artifact_id}/tags` | `get_artifact_tags` | Read an Artifact's current tags | +| `PUT` | `/v1/scopes/{scope_id}/artifacts/{family}/{artifact_id}/tags` | `replace_artifact_tags` | Replace an Artifact's complete tag set | +| `GET` | `/v1/scopes/{scope_id}/artifacts/memory/{artifact_id}/entries/{entry_id}/tags` | `get_memory_entry_tags` | Read a Memory entry's current tags | +| `PUT` | `/v1/scopes/{scope_id}/artifacts/memory/{artifact_id}/entries/{entry_id}/tags` | `replace_memory_entry_tags` | Replace a Memory entry's complete tag set | +| `POST` | `/v1/scopes/{scope_id}/artifact-tags/query` | `query_artifact_tags` | Retrieve visible current targets by exact tags | + +The two target shapes expose the same `ArtifactTagSet` schema, validation, authorization rules, and repository. The +extra `entries/{entry_id}` path segment expresses containment; it does not create a second tag model or table. Paths +name current logical targets only. Exact Artifact Revision and Memory entry-version paths have no tag subresource. + +### Get + +Both GET operations return `200 ArtifactTagSet` and an opaque `ETag`. They support `If-None-Match` and return +`304 Not Modified` without a body when it matches. A visible target with no assignments returns `200`, an empty set, +and an ETag. A missing or non-visible target follows the same `404` behavior as reading that target. + +### Replace + +Both PUT operations accept the complete replacement set: + +```json +{ + "tags": ["Customer-A", "cockpit", "verified"] +} +``` + +`If-Match` is required. The server resolves the opaque validator to the expected target-bound tag state, performs the +repository replacement, and returns `200 ArtifactTagSet` plus the new ETag. Missing `If-Match` returns +`428 Precondition Required`; a mismatch returns `412 Precondition Failed`. The response does not disclose the current +ETag or tag values to a caller that cannot read the target. PUT does not create an Artifact Revision. + +### Query + +```json +{ + "tags": ["customer-a", "verified"], + "match": "all", + "families": ["memory", "experience", "skill"], + "target_types": ["artifact", "memory_entry"], + "include_inactive": false, + "limit": 50, + "cursor": null +} +``` + +`tags` has between 1 and 16 unique normalized values. `match` defaults to `all`. Omitted `families` and +`target_types` select every supported value. `include_inactive` defaults to false and never bypasses authorization; +it only expands lifecycle selection for callers already allowed to inspect inactive content. + +Each page item contains the logical `target`, all current tags, and one current exact content reference: + +```json +{ + "target": { + "type": "memory_entry", + "family": "memory", + "artifact_id": "memory", + "entry_id": "mem_ent_123" + }, + "current": { + "memory_ref": { + "family": "memory", + "artifact_id": "memory", + "revision": 12 + }, + "entry_id": "mem_ent_123", + "entry_version_id": "mem_ver_456" + }, + "tags": ["Customer-A", "cockpit", "verified"] +} +``` + +Artifact targets use an `ArtifactReference` in `current`; Memory entry targets use `MemoryCitation`. Items are ordered +by `(family, target_type, artifact_id, target_id)` using UTF-8 byte order. The opaque cursor binds the Scope, normalized +tag keys, match mode, selected families, target types, lifecycle selection, caller, expiration, and last ordering key. +An invalid or filter-mismatched cursor returns `400 Bad Request`; an expired cursor returns `410 Gone`. + +Tag mutations between pages can change membership. Pagination guarantees deterministic keyset traversal for each +query but does not claim a database snapshot across requests. + +## Existing list and search integration + +The following existing request surfaces gain optional filtering with the same normalized `tags` and `match` +semantics: + +- RFC 1437's `GET /v1/scopes/{scope_id}/artifacts/{family}` adds repeatable `tag` and optional + `tag_match=all|any` query parameters; +- Memory entry listing adds an optional `tag_filter` request field; and +- Memory search adds an optional `tag_filter` request field. + +The parameters or field are absent by default and therefore preserve existing behavior. A present filter must contain +at least one tag. `tag_match` without `tag` is invalid. Artifact-head listing only matches `artifact` targets. +Memory-entry list and search only match `memory_entry` targets. Artifact-list cursors additionally bind the normalized +tags and match mode; the RFC 1437 invalid, mismatched, and expired cursor statuses remain unchanged. + +Existing Artifact-list, Memory-entry-list, and Memory-search item schemas do not gain a `tags` field; their tag filter +changes eligibility only. The dedicated tag query includes current tags, and callers can GET the logical target's tag +subresource when current catalog metadata is required. Exact historical Artifact Revision and Memory entry-version +responses do not acquire a `tags` field. + +Memory search applies tag eligibility inside both FTS and vector candidate queries before channel limits. Hybrid +search applies the same eligible target set to both channels before fusion and reranking. A backend that cannot apply +the filter before top-k must report the combined mode unavailable rather than silently over-fetching and returning an +incomplete result. + +This RFC does not add tag constraints to automatic `PreparedContext` assembly. A later use case may add a typed +selection profile, but tags never enter a model prompt merely because they exist. + +## Authorization and trust boundary + +`scope_id` is a business partition, not proof of authority. Tag operations use the same Server authentication and +authorization boundary as the target resource: + +- reading tags requires permission to read the target; +- replacing tags requires permission to mutate the target's catalog metadata; +- query results include only targets the principal may discover and read; and +- `include_inactive` does not broaden resource access. + +The initial implementation may map metadata mutation to the existing target write authority. A deployment must not +infer access from tag values, create grants from tags, or use tags as a substitute for the access-control Resource +Profile. If a later product needs delegated taxonomy management without content write authority, that requires a +separate action and audit design. + +Tags are untrusted display strings. Dashboard rendering must escape them and must not interpret them as HTML, +Markdown, URLs, commands, or CSS classes. Search and application code must use bound parameters. Tags are never +executed and are never injected into Agent instructions by this RFC. + +## Publication, import, and lineage + +Tag assignments are Scope-local catalog state and do not participate in: + +- `ArtifactLineage`; +- content or package digests; +- publication digests; +- Source evidence; +- Candidate approval; or +- managed Skill package metadata. + +Publishing, copying, importing, or forking an Artifact does not copy assignments. Destination assignment is a separate +authorized write. Source tags may be displayed to an authorized publisher before that write, but they are not treated +as content provenance. + +## Compatibility and migration + +The relational initializer adds `pc_artifact_tags` for every supported database profile. Existing Artifacts and Memory +entries require no backfill and behave as if they have an empty tag set. + +No migration copies values from `ContentSource.metadata`, `SkillContent.metadata`, Memory `kind`, lifecycle state, +review status, or integration provenance. Those fields have different authorities and semantics. + +All additions to existing list and search requests are optional. Older clients that send neither the Artifact-list +parameters (`tag` and `tag_match`) nor the Memory request field (`tag_filter`) retain their current result semantics. +The new operations and schemas are added to `openapi/powercontext.yaml`; generated Python and integration contracts +must be regenerated through the repository's normal contract workflow. + +## Observability + +The Server may record operation outcome, target type, family, submitted tag count, filter tag count, match mode, result +count, and latency. Logs, traces, metrics, and error messages must not record raw tag values. A normalized tag digest +may be used for correlation only when deployment policy permits it. + +Tag mutation must use the ordinary authenticated audit boundary when available. This RFC does not add a historical +assignment table; a deployment that requires a complete tag-change ledger must keep the feature disabled or add the +ledger through a follow-up design before claiming that guarantee. + +## Delivery plan + +The implementation is split into two reviewable vertical slices: + +1. Add target models, normalization, the shared table and repository, ETag-guarded read/replace plus query, + OpenAPI-generated contracts, and deterministic SQLite and OceanBase/seekDB behavior tests. +2. Add current Artifact and Memory-entry list filters, pre-top-k Memory search filtering, and the minimal Dashboard tag + editor and exact filter. + +The feature is not complete after only adding the table. The first customer-visible release requires both slices so a +user can assign a tag, observe it, retrieve the target with it, and remove it without editing Artifact content. + +## Acceptance criteria + +The RFC is implemented only when all of the following observable scenarios pass: + +1. A caller assigns and reads tags on Experience, Skill, Handoff, and whole-Memory Artifact targets through the same + tag-set semantics. +2. A caller assigns a different set to one Memory entry without changing tags on another entry in the same Memory. +3. Revising an Artifact or Memory entry preserves its logical target tags and changes no tag because of Revision alone. +4. Replacing tags changes no Artifact Revision, content digest, lineage, Memory entry version, FTS text, or embedding. +5. Empty-set replacement removes all assignments and remains retrievable as an empty tag set. +6. `all` and `any` matching return correct targets across families, with duplicate-normalized input rejected. +7. Artifact listing and Memory entry listing preserve existing lifecycle defaults while applying tag filters. +8. FTS, vector, and hybrid Memory search apply tag eligibility before their candidate limits and return no untagged + result. +9. A stale tag-set ETag cannot overwrite a concurrent replacement, including when the initial tag set was empty. +10. Pagination rejects a cursor reused with different normalized filters and returns deterministic keyset order. +11. A principal cannot discover, read, or mutate another target's tags without the corresponding target authority. +12. Publication to another Scope creates no destination assignments and cannot expose raw source tags implicitly. +13. Raw tag values do not appear in telemetry, and Dashboard rendering treats adversarial values as inert text. +14. Existing unfiltered API behavior and exact historical content responses remain unchanged. +15. Schema, repository, HTTP contract, generated-client, and supported backend tests pass through repository-standard + commands. + +# Drawbacks + +A polymorphic assignment table cannot use one ordinary conditional foreign key to validate both Artifact and Memory +entry targets. The repository must enforce Memory-entry existence transactionally. A universal catalog-item registry +would provide a single foreign key, but it would add another persistent identity layer and another table before any +other feature needs one. + +Current logical tags are not reconstructible at an earlier time. Exact Artifact content remains reproducible, but the +tag set is only current catalog state. Deployments requiring historical taxonomy audit need an additional event or +history design. + +The two query indexes increase write and storage cost. The bounded per-target tag count keeps this cost predictable, +and tag writes are expected to be much less frequent than reads. + +Tag pre-filtering must be implemented in every supported Memory search backend. This is more work than filtering final +hits, but it is required for correct top-k behavior. + +# Rationale and alternatives + +## Store tags in Artifact content + +This gives immutable historical tags but turns a classification edit into a content Revision. It changes content +digests, lineage expectations, CAS behavior, and derived indexes without changing reusable knowledge. It is rejected +for user-managed catalog tags. + +## Add a JSON tag column to `pc_artifact_heads` + +This is a small schema change for whole Artifacts, but it cannot distinguish individual entries in the standard +one-Memory-per-Scope model. Portable indexed `all`/`any` queries also differ across SQLite and MySQL-compatible +backends. It is rejected. + +## Add one table for Artifacts and another for Memory entries + +This provides direct foreign keys and simple family-local joins. It fragments cross-family queries and duplicates tag +normalization, mutation, pagination, and API behavior. The single polymorphic table keeps one assignment contract and +accepts one explicit application-level Memory-entry invariant. + +## Add a universal catalog-item registry + +A registry could give every nested and top-level resource a uniform ID and let tags reference one parent table. It +would require new lifecycle, migration, ownership, deletion, and synchronization semantics for all current resources. +Tags alone do not justify that abstraction in the first delivery. + +## Rebuild Memory so every entry is an Artifact + +This would make physical tag targets uniform, but it would replace the existing Memory Manifest, atomic collection +Revision, entry-version, citation, flush, and index model. It is disproportionate to the customer requirement and is +rejected. + +## Use tags as arbitrary metadata keys and values + +The requested behavior is set membership and exact filtering. A generic nested metadata object requires type, +operator, indexing, conflict, and authorization semantics that are not needed here. A later key/value metadata feature +must not silently reinterpret tags. + +## Do nothing + +Users would continue encoding classifications in content, keeping external spreadsheets, or relying on ambiguous text +queries. None provides a consistent, Scope-aware, cross-family retrieval contract. + +# Prior art + +PowerContext already separates immutable Artifact Revisions from mutable current heads and rebuildable search +projections. RFC 0014 defines Memory entry identity and exact citation; RFC 0019 defines one current Memory Artifact per +Scope in the standard profile; RFC 0048 defines Handoff as a self-contained Artifact lifecycle; and RFC 0051 defines +Experience and managed Skill as independent Artifact families. This RFC applies the same separation of logical identity +and immutable content to user-managed classification. + +RFC 1396 separates resource authorization from Artifact content and warns that Scope identity is not authorization. +Tags follow that boundary: they may help a user find a resource but never decide whether the user may access it. + +RFC 1437 establishes the Scope-owned Artifact URI tree, opaque HTTP validators for mutable current representations, +and caller- and query-bound expiring cursors. This RFC extends that tree with logical-target tag subresources and +extends Artifact listing without changing exact Revision responses or treating `tag_digest` as an Artifact ETag. + +No external system is normative for this RFC. Common repository and issue trackers demonstrate that mutable labels +can organize immutable or versioned content, but PowerContext's Memory-entry containment and exact citation model +require the target contract defined here. + +# Unresolved questions + +No unresolved question blocks acceptance of the first-delivery contract. + +The implementation PR must still confirm backend-specific query plans for the bounded `all` filter and select exact +index names that fit repository naming limits. These are implementation validation details and must not change public +normalization, matching, ordering, or pre-top-k filtering semantics. + +The following questions are intentionally outside this RFC: + +- whether organizations need managed tag definitions, colors, descriptions, aliases, or rename operations; +- whether tag mutation needs a separately delegable authorization action; +- whether some publication workflow should explicitly offer to copy selected tags; +- whether automated classification can safely propose, but not silently assign, tags; and +- whether a complete historical tag-assignment ledger is required. + +# Future possibilities + +A later RFC may add a Scope-local tag catalog with descriptions, display color, aliases, usage counts, controlled +rename, or delegated taxonomy management. Such a catalog would describe tags; `pc_artifact_tags` would remain the +assignment relation. + +Another extension may let users save named search views that combine tags, families, lifecycle state, and text query. +Saved views must remain queries rather than authorization policy. + +Model-assisted classification may propose tags through a reviewed Candidate-like flow. Models must not assign tags +silently, and proposed tags must remain untrusted until an authorized user accepts them. + +If multiple nested resource types need tags, access control, favorites, comments, and other catalog metadata, the +project may then justify a universal logical-resource registry. That decision should be based on several proven uses, +not introduced speculatively by this RFC. diff --git a/docs/zh/rfcs/0000_artifact_tags.md b/docs/zh/rfcs/0000_artifact_tags.md new file mode 100644 index 000000000..bedbfc4ed --- /dev/null +++ b/docs/zh/rfcs/0000_artifact_tags.md @@ -0,0 +1,646 @@ +- Proposal Name: `unified_artifact_tags` +- Start Date: 2026-09-05 +- RFC PR: [oceanbase/powercontext#0000](https://github.com/oceanbase/powercontext/pull/0000) +- Related Discussion: [oceanbase/powercontext#1466](https://github.com/oceanbase/powercontext/issues/1466) +- Related RFCs: [RFC 0014](0014_memory_layer_design.md)、[RFC 0019](0019_local_source_memory_runtime.md)、 + [RFC 0048](0048_handoff_artifact.md)、[RFC 0051](0051_experience_skill_artifact_families.md)、 + [RFC 1396](1396_handoff_access_control.md) 和 [RFC 1437](1437_source_artifact_rest_api.md) + +# Summary + +本 RFC 为 PowerContext 管理的内容增加 Scope-local 自定义标签。用户可以给 managed Experience、Skill、Handoff、 +整个 Memory Artifact 或单条 Memory entry 添加 `customer-a`、`cockpit`、`verified` 等标签,再按精确标签成员关系 +检索当前资源。 + +该能力为所有受支持 target 共用一张 `pc_artifact_tags` 分配表。每一行都保留所属 Artifact 身份,并标识该 Artifact +自身或其中的一条逻辑 Memory entry。这样既能提供统一的产品与 API 模型,也不会把 Memory entry 伪装成独立 +Artifact。 + +标签是逻辑 target 的可变 catalog 属性,不是 Artifact content、Source metadata、lineage、授权策略、生命周期状态或 +Prompt 指令。修改标签不会创建 Artifact Revision、改变 content digest、重建 embedding 或改变精确 citation。标签留在 +owner Scope 内;除非后续 contract 显式要求,否则 Artifact publication 不复制标签。 + +第一阶段交付标签读取和 compare-and-swap 替换、精确的 `all` 与 `any` 标签查询、current-resource list 和 Memory +search 的可选标签过滤,以及最小 Dashboard 编辑器和筛选器。标签层级、颜色、别名、自动打标、基于标签的授权及历史 +标签快照不进入第一阶段。 + +# Motivation + +团队会在同一个 Scope 中积累大量 Memory entry、Experience 和 Skill。文本与语义搜索能回答“哪些内容与查询相似”, +但不能稳定回答以下组织问题: + +- 哪些条目属于客户 A? +- 哪些 Skill 已在智能座舱环境中验证? +- 哪些 Experience 描述发布操作,而不是推理行为? +- 哪些已退役或 inactive 资源仍属于某个合规检查集合? + +用户可以把这些信息编码进 content,但这样会混淆分类与知识本身。修改分类将产生新的不可变 Revision、改变 content +digest,并可能重建 search projection,即使可复用内容没有变化。 + +现有名为 `metadata` 的字段不能提供共享方案。`ContentSource.metadata` 包含 provenance 和 `kind` 等影响行为的字段; +managed Skill metadata 属于精确 Skill content,并参与其 search projection。把任一字段当成通用可变标签袋,都会模糊 +ownership、version、indexing 和 security 语义。 + +持久化模型还有两种用户可见的 target 粒度: + +- Experience、Skill、Handoff 和整个 Memory 是由 `artifact_id` 选择的逻辑 Artifact lifecycle; +- 用户看到的一条事实或偏好 Memory,是 Scope 的 Memory Artifact 中由 `entry_id` 选择的逻辑 entry。 + +只在 `pc_artifact_heads` 增加一个 JSON 标签列,会让标准 one-Memory-per-Scope profile 中的所有 entry 共用同一组标签。 +分别增加 family-specific 标签存储虽然能保留粒度,却会分裂 API 和跨 Family 查询路径。PowerContext 需要一个能处理两种 +target 形状的显式分配模型。 + +# Guide-level explanation + +## 用户模型 + +标签是用户在一个 Scope 内为一个 current logical target 编写的字符串,用于精确分组和筛选。它不会证明某个断言为真, +不会批准 Artifact、授予访问权或指示 Agent。 + +所有受支持资源具有相同的可见行为: + +```text +Scope: vehicle-assistant + +Memory entry: "驾驶员冬季偏好 24 C" +Tags: [customer-a, cockpit, preference] + +Experience: "修改 OpenAPI 后重新生成 Client" +Tags: [release, verified] + +Skill: "vehicle-log-triage" +Tags: [customer-a, diagnostics] + +Handoff: "完成座舱延迟调查" +Tags: [customer-a, in-progress] +``` + +当分类作用于整个集合而不是单条 entry 时,用户也可以给整个 Memory Artifact 打标签。UI 必须区分 `Memory` 与 +`Memory entry`,避免把集合标签误认为条目标签。 + +## 添加和删除标签 + +Dashboard 将标签显示为 inert、已转义的 chip。具有 write authority 的用户打开资源,编辑完整标签集并保存。保存操作 +以编辑器读取到的标签集为条件,避免两个编辑器静默覆盖对方。 + +例如,读取当前集合时,response header 返回 opaque `ETag`,body 为: + +```json +{ + "tags": ["cockpit", "customer-a"], + "tag_digest": "sha256:92f..." +} +``` + +用户将其替换为: + +```json +{ + "tags": ["cockpit", "customer-a", "verified"] +} +``` + +client 在 `If-Match` 中发送读取到的 ETag,而不是根据 `tag_digest` 构造 precondition。 + +用相同 normalized tags 替换是幂等操作。用空数组替换会删除全部标签,但不会删除、退役或修订 target。 + +## 按标签检索 + +标签经过规范化后按精确值匹配。筛选 `customer-a` 不会匹配 `customer-a-archive`;名为 `release/security` 的标签也 +不会隐含一个名为 `release` 的父标签。 + +多个标签具有显式 match mode: + +- `all` 选择拥有所有请求标签的 target; +- `any` 选择至少拥有一个请求标签的 target。 + +例如: + +```json +{ + "families": ["memory", "experience", "skill"], + "target_types": ["artifact", "memory_entry"], + "tags": ["customer-a", "verified"], + "match": "all", + "limit": 50 +} +``` + +响应返回 logical target 及其当前精确 reference。Artifact 结果包含当前 `ArtifactReference`,Memory entry 结果包含当前 +`MemoryCitation`。精确 reference 让调用方读取内容时不需要再次解析 `latest`。 + +文本检索可以组合 query 与标签过滤。PowerContext 在 FTS、vector top-k selection、fusion 或 reranking 前,对 eligible +candidate set 应用标签过滤。先截断 top-k 再过滤是错误的,因为相关且带标签的条目可能在过滤发生前就已被排除。 + +## Revision 行为 + +标签跟随逻辑身份: + +```text +Experience exp-1 Revision 1 --\ +Experience exp-1 Revision 2 ----> logical exp-1 的 tags + +Memory entry entry-1 Version 1 --\ +Memory entry entry-1 Version 2 ----> logical entry-1 的 tags +``` + +因此,修订 `exp-1` 或 `entry-1` 会保留其标签;给任一 target 打标签都不会创建 content Revision。精确历史 Artifact +或 Memory citation 仍然只声明不可变 content 与 evidence,不会隐式获得历史标签快照。 + +## Scope 与 publication 行为 + +标签属于分配它们的 Scope。将 Artifact 发布或复制到另一个 Scope 时,目标 logical Artifact 以空标签集创建。该默认行为 +避免泄漏客户名称、内部工作流分类或其他本地分类。调用方可以在 publication 后显式分配目标标签。 + +# Reference-level explanation + +## Goals + +第一阶段实现必须: + +- 为 managed Artifact 与 logical Memory entry 提供一个标签模型; +- 用一张表持久化所有分配关系; +- 使标签变更独立于不可变 Artifact 和 Memory-entry content; +- 以确定性规范化支持精确 `all` 与 `any` 过滤; +- 正确组合标签与 current Artifact listing、Memory retrieval; +- 保留 Scope isolation、target visibility 和现有 lifecycle filter; +- 提供并发安全且幂等的 complete-set replacement; +- 暴露足够的 current exact identity,使查询结果可以被安全解析。 + +## Non-goals + +第一阶段不定义: + +- 标签颜色、描述、别名、层级、继承或标签定义 catalog; +- 由模型、Source metadata 或 content keyword 自动生成标签; +- 把标签作为授权、批准、信任、生命周期、路由或保留策略; +- 历史标签快照或标签分配 audit history; +- 通过 lineage、publication、fork、import 或 Handoff evidence 传播标签; +- revision-specific tag; +- 允许 Agent 自动修改标签的 MCP tool; +- 任意 key/value Artifact metadata。 + +## 术语与 target identity + +`ArtifactTagTarget` 是一个 discriminated union: + +```text +ArtifactTagTarget = + ArtifactTarget { + type: "artifact", + family: string, + artifact_id: string + } + | MemoryEntryTarget { + type: "memory_entry", + family: "memory", + artifact_id: string, + entry_id: string + } +``` + +`ArtifactTarget` 标识一个 logical Artifact lifecycle,并刻意省略 `revision`。`MemoryEntryTarget` 标识一条 logical +Memory entry,并刻意省略 Memory Revision 与 `entry_version_id`。所属 `artifact_id` 保持显式,因为 `entry_id` 的作用域 +是其 Memory Artifact。 + +canonical persistence form 使用 `target_id`: + +| Target | `family` | `artifact_id` | `target_type` | `target_id` | +| --- | --- | --- | --- | --- | +| Experience | `experience` | Experience ID | `artifact` | 相同 Experience ID | +| Skill | `skill` | Skill ID | `artifact` | 相同 Skill ID | +| Handoff | `handoff` | Handoff ID | `artifact` | 相同 Handoff ID | +| 整个 Memory | `memory` | Memory ID | `artifact` | 相同 Memory ID | +| Memory entry | `memory` | 所属 Memory ID | `memory_entry` | entry ID | + +新的 nested resource kind 不能复用 `memory_entry` 或重载 `target_id`。它们需要在后续 contract 中定义显式 target type +和 validation rule。 + +## 标签值与规范化 + +提交的标签必须满足以下全部条件: + +- 是 1 至 64 个 Unicode code point 的字符串; +- 没有前导或尾随空白; +- 不包含 Unicode control、surrogate 或 unassigned code point; +- normalized key 不超过 128 个 Unicode code point; +- 完整提交集合最多包含 32 个标签。 + +PowerContext 将提交值保留为用于展示的 `tag`,先应用 Unicode NFC,再应用 Unicode default case folding,生成 +`tag_key`。它不会折叠内部空白、拆分标点、解析 `/`、翻译、词干化或推断层级。`tag` 与 `tag_key` 都在规范化后校验。 + +两个拥有相同 `tag_key` 的提交标签视为重复,整个请求会被拒绝;Server 不会静默选择某种展示拼写。后续成功替换可以只 +改变保留的展示拼写,同时保持相同 `tag_key`。 + +标签按 `tag_key` 的 UTF-8 byte 升序排列。响应和 digest 计算都使用该顺序,避免数据库 collation 改变公开行为。 + +## Persistence + +共享关系型 schema 只新增一张业务表: + +```text +pc_artifact_tags + scope_id identity string, not null + family identity string, not null + artifact_id identity string, not null + target_type identity string, not null + target_id identity string, not null + tag_key identity string, not null + tag display string, not null + assigned_at UTC timestamp, not null + + primary key ( + scope_id, + family, + artifact_id, + target_type, + target_id, + tag_key + ) + + foreign key (scope_id, family, artifact_id) + references pc_artifact_heads (scope_id, family, artifact_id) + on delete cascade + + check target_type in ('artifact', 'memory_entry') + check target_type != 'artifact' or target_id = artifact_id + check target_type != 'memory_entry' or family = 'memory' +``` + +该表包含以下 secondary indexes: + +```text +(scope_id, family, tag_key, target_type, artifact_id, target_id) +(scope_id, tag_key, family, target_type, artifact_id, target_id) +``` + +primary key 支持加载一个 target 的标签;第一个 secondary index 支持 family-specific 过滤,第二个支持 Scope 内跨 Family +查询。实现必须使用 binary identity comparison 或应用构造的 `tag_key`,不能依赖数据库默认大小写或 locale collation。 + +一张表无法用普通 conditional foreign key 同时校验 `pc_artifact_heads` 和 `pc_memory_entry_heads`。每一行都强制所属 +Artifact foreign key。对于 `memory_entry`,repository 还必须在同一个 transaction 中锁定并校验当前 +`(scope_id, memory_artifact_id, entry_id)` head,再修改分配关系;缺少 entry 时 transaction 拒绝请求。这是明确的 +application invariant,不是 best-effort cleanup rule。 + +inactive Memory entry 与 deprecated 或 retired Artifact 保留其分配关系。标签查询先应用现有 visibility 和 lifecycle +selection,再返回结果。拥有 target write authority 的调用方可以重新组织标签,而不会重新激活或修订 content。 + +## 标签集与 digest + +`ArtifactTagSet` 包含: + +```json +{ + "scope_id": "vehicle-assistant", + "target": { + "type": "memory_entry", + "family": "memory", + "artifact_id": "memory", + "entry_id": "mem_ent_123" + }, + "tags": ["cockpit", "customer-a"], + "tag_digest": "sha256:..." +} +``` + +`tag_digest` 是对象 `{"tags": [...]}` 的 RFC 8785 canonical JSON 的 SHA-256 digest;标签按 canonical `tag_key` +顺序排列,每个 array item 使用保留的展示字符串。空集合也有稳定 digest。它是 tag set 的 content checksum,不是 +client 可见的 compare-and-swap token、Artifact content digest、Memory entry content hash 或 authorization generation。 + +HTTP operation 使用绑定完整 logical target identity 与 `tag_digest` 的 opaque ETag。client 不能假定 ETag 等于、包含或 +可以由 `tag_digest` 重建。 + +## Repository contract 与 transaction + +tag repository 暴露三个操作: + +```text +get(scope_id, target) -> ArtifactTagSet + +replace( + scope_id, + target, + expected_tag_digest, + tags +) -> ArtifactTagSet + +query( + scope_id, + tags, + match, + families, + target_types, + lifecycle_selection, + limit, + cursor +) -> ArtifactTagPage +``` + +`replace` 在一个 transaction 内执行以下步骤: + +1. 解析并授权 target,同时不返回隐藏的存在性细节。 +2. 锁定所属 Artifact head;对于 Memory entry,还要锁定并校验其 current entry head。 +3. 加载完整 current tag set 并计算 digest。 +4. expected digest 不匹配时,按 precondition failure 拒绝请求。 +5. 校验并规范化完整 replacement set。 +6. 删除 replacement 中缺少的分配,并插入或更新保留的展示值。 +7. 返回 canonical set 与新 digest。 + +即使 current tag set 为空,锁定已有所属 head 也能串行化 replacement。实现不能依赖锁定零行 assignment,也不能依赖 +process-local lock,因为二者都不能提供所需的分布式 compare-and-swap 行为。 + +如果当前 canonical set 与请求集合相同,`replace` 幂等成功,并且不修改任何 row 或 `assigned_at` 值。 + +## HTTP contract + +第一阶段 contract 在 RFC 1437 建立的 Scope resource tree 下新增五个 operation: + +| Method | Path | operationId | Purpose | +| --- | --- | --- | --- | +| `GET` | `/v1/scopes/{scope_id}/artifacts/{family}/{artifact_id}/tags` | `get_artifact_tags` | 读取 Artifact current tags | +| `PUT` | `/v1/scopes/{scope_id}/artifacts/{family}/{artifact_id}/tags` | `replace_artifact_tags` | 替换 Artifact 完整 tag set | +| `GET` | `/v1/scopes/{scope_id}/artifacts/memory/{artifact_id}/entries/{entry_id}/tags` | `get_memory_entry_tags` | 读取 Memory entry current tags | +| `PUT` | `/v1/scopes/{scope_id}/artifacts/memory/{artifact_id}/entries/{entry_id}/tags` | `replace_memory_entry_tags` | 替换 Memory entry 完整 tag set | +| `POST` | `/v1/scopes/{scope_id}/artifact-tags/query` | `query_artifact_tags` | 按精确 tag 检索可见 current targets | + +两种 target 形状使用相同的 `ArtifactTagSet` schema、validation、authorization rule 与 repository。额外的 +`entries/{entry_id}` path segment 只表达 containment,不创建第二套标签模型或表。path 只命名 current logical target; +精确 Artifact Revision 与 Memory entry-version path 不提供 tag subresource。 + +### Get + +两个 GET operation 都返回 `200 ArtifactTagSet` 和 opaque `ETag`。它们支持 `If-None-Match`,匹配时返回没有 body 的 +`304 Not Modified`。没有任何 assignment 的可见 target 返回 `200`、空集合和 ETag。缺失或不可见的 target 与读取该 +target 一样返回 `404`。 + +### Replace + +两个 PUT operation 都接收完整 replacement set: + +```json +{ + "tags": ["Customer-A", "cockpit", "verified"] +} +``` + +`If-Match` 是必需 header。server 将 opaque validator 解析为预期的 target-bound tag state,执行 repository replacement, +然后返回 `200 ArtifactTagSet` 和新 ETag。缺少 `If-Match` 返回 `428 Precondition Required`;不匹配返回 +`412 Precondition Failed`。response 不向不能读取 target 的调用方泄露 current ETag 或 tag value。PUT 不创建 +Artifact Revision。 + +### Query + +```json +{ + "tags": ["customer-a", "verified"], + "match": "all", + "families": ["memory", "experience", "skill"], + "target_types": ["artifact", "memory_entry"], + "include_inactive": false, + "limit": 50, + "cursor": null +} +``` + +`tags` 包含 1 至 16 个 normalized 后唯一的值。`match` 默认为 `all`。省略 `families` 和 `target_types` 表示选择 +全部支持值。`include_inactive` 默认为 false,并且永远不会绕过授权;它只为已经能够 inspect inactive content 的调用方 +扩展 lifecycle selection。 + +每个 page item 包含 logical `target`、所有 current tags,以及一个 current exact content reference: + +```json +{ + "target": { + "type": "memory_entry", + "family": "memory", + "artifact_id": "memory", + "entry_id": "mem_ent_123" + }, + "current": { + "memory_ref": { + "family": "memory", + "artifact_id": "memory", + "revision": 12 + }, + "entry_id": "mem_ent_123", + "entry_version_id": "mem_ver_456" + }, + "tags": ["Customer-A", "cockpit", "verified"] +} +``` + +Artifact target 的 `current` 使用 `ArtifactReference`;Memory entry target 使用 `MemoryCitation`。item 按 +`(family, target_type, artifact_id, target_id)` 的 UTF-8 byte order 排序。opaque cursor 绑定 Scope、normalized tag +key、match mode、所选 family、target type、lifecycle selection、caller、expiration 和最后一个 ordering key。非法或 +filter 不匹配的 cursor 返回 `400 Bad Request`;过期 cursor 返回 `410 Gone`。 + +page 之间发生的标签变更可能改变成员关系。pagination 为每次查询提供确定性 keyset traversal,但不承诺跨请求数据库 +snapshot。 + +## 现有 list 与 search 集成 + +以下现有 request surface 增加可选 filter,使用相同的 normalized `tags` 和 `match` 语义: + +- RFC 1437 的 `GET /v1/scopes/{scope_id}/artifacts/{family}` 增加可重复的 `tag` query parameter 和可选的 + `tag_match=all|any`; +- Memory entry listing 增加可选 `tag_filter` request field; +- Memory search 增加可选 `tag_filter` request field。 + +parameter 或 field 默认缺失,因此保留现有行为。filter 存在时必须至少包含一个标签;没有 `tag` 时设置 `tag_match` 属于 +非法请求。Artifact-head listing 只匹配 `artifact` target;Memory-entry list 和 search 只匹配 `memory_entry` target。 +Artifact-list cursor 还要绑定 normalized tag 与 match mode;RFC 1437 的非法、不匹配和过期 cursor status 保持不变。 + +现有 Artifact-list、Memory-entry-list 和 Memory-search item schema 不新增 `tags` 字段;tag filter 只改变 eligibility。 +专用 tag query 包含 current tags;需要 current catalog metadata 时,调用方可以 GET logical target 的 tag subresource。 +精确历史 Artifact Revision 和 Memory entry-version response 不新增 `tags` 字段。 + +Memory search 在两个 FTS 和 vector candidate query 内、channel limit 前应用 tag eligibility。Hybrid search 在 fusion 与 +reranking 前,对两个 channel 应用同一个 eligible target set。无法在 top-k 前应用过滤的 backend 必须报告 combined mode +不可用,不能静默 over-fetch 并返回不完整结果。 + +本 RFC 不给自动 `PreparedContext` assembly 增加标签约束。后续用例可以新增 typed selection profile,但标签不会因为 +存在而进入 model prompt。 + +## Authorization 与 trust boundary + +`scope_id` 是业务分区,不是 authority 证明。标签 operation 使用与 target resource 相同的 Server authentication 和 +authorization boundary: + +- 读取标签需要 target read 权限; +- 替换标签需要修改 target catalog metadata 的权限; +- query result 只包含 principal 可以发现和读取的 target; +- `include_inactive` 不会扩大资源访问范围。 + +初始实现可以把 metadata mutation 映射到现有 target write authority。deployment 不能从标签值推导访问权、用标签创建 +grant,或用标签代替 Access Control Resource Profile。如果后续产品需要在没有 content write authority 的情况下委派 +taxonomy management,则需要独立 action 与 audit 设计。 + +标签是不可信的 display string。Dashboard rendering 必须转义标签,不能把它解释为 HTML、Markdown、URL、command 或 +CSS class。Search 与 application code 必须使用 bound parameter。本 RFC 永远不会执行标签,也不会把标签注入 Agent +instruction。 + +## Publication、import 与 lineage + +标签分配是 Scope-local catalog state,不参与: + +- `ArtifactLineage`; +- content 或 package digest; +- publication digest; +- Source evidence; +- Candidate approval; +- managed Skill package metadata。 + +发布、复制、导入或 fork Artifact 都不会复制 assignment。授权 publisher 可以在写入前看到 source tag,但不能把它们当成 +content provenance。目标分配是独立的授权写入。 + +## Compatibility 与 migration + +关系型 initializer 为每个受支持 database profile 新增 `pc_artifact_tags`。现有 Artifact 和 Memory entry 不需要 backfill, +行为等同于拥有空标签集。 + +任何 migration 都不会从 `ContentSource.metadata`、`SkillContent.metadata`、Memory `kind`、lifecycle state、review +status 或 integration provenance 复制值。这些字段具有不同的 authority 和语义。 + +现有 list 与 search request 的所有新增字段都是 optional。既未发送 Artifact-list parameter(`tag` 与 +`tag_match`),也未发送 Memory request field(`tag_filter`)的旧 Client 保持 current result 语义。新 operation 与 +schema 添加到 `openapi/powercontext.yaml`;generated Python 和 integration contract 必须通过仓库正常 contract +workflow 重新生成。 + +## Observability + +Server 可以记录 operation outcome、target type、family、提交标签数量、过滤标签数量、match mode、result count 和 +latency。log、trace、metric 与 error message 不能记录 raw tag value。只有 deployment policy 允许时,才能用 normalized +tag digest 进行关联。 + +当普通 authenticated audit boundary 可用时,标签 mutation 必须使用该边界。本 RFC 不增加 historical assignment +table;需要完整标签变更账本的 deployment 必须禁用该能力,或在声称拥有该保证前通过后续设计增加账本。 + +## Delivery plan + +实现分为两个可评审的 vertical slice: + +1. 增加 target model、normalization、共享表和 repository、ETag-guarded read/replace 与 query、OpenAPI generated + contract,以及 SQLite 与 OceanBase/seekDB 的确定性行为测试。 +2. 增加 current Artifact 与 Memory-entry list filter、pre-top-k Memory search filtering,以及最小 Dashboard 标签 + editor 与精确筛选。 + +只增加表不代表功能完成。首个 customer-visible release 需要两个 slice 都完成,使用户能够分配标签、看到标签、通过标签 +检索 target,并在不修改 Artifact content 的情况下移除标签。 + +## Acceptance criteria + +只有以下可观察场景全部通过,才视为本 RFC 已实现: + +1. 调用方通过相同 tag-set 语义给 Experience、Skill、Handoff 和 whole-Memory Artifact target 分配和读取标签。 +2. 调用方给一条 Memory entry 分配独立标签集,不改变同一 Memory 中其他 entry 的标签。 +3. 修订 Artifact 或 Memory entry 会保留其 logical target tag,单纯 Revision 变化不会改变任何标签。 +4. 替换标签不会改变 Artifact Revision、content digest、lineage、Memory entry version、FTS text 或 embedding。 +5. empty-set replacement 删除所有 assignment,随后仍可以读取为空 tag set。 +6. `all` 和 `any` matching 跨 Family 返回正确 target,并拒绝 normalized 后重复的输入。 +7. Artifact listing 和 Memory entry listing 在应用标签过滤时保留现有 lifecycle default。 +8. FTS、vector 和 hybrid Memory search 在各自 candidate limit 前应用 tag eligibility,不返回未标记结果。 +9. 即使初始标签集为空,过期 tag-set ETag 也不能覆盖并发 replacement。 +10. pagination 拒绝用不同 normalized filter 重用 cursor,并返回确定性 keyset order。 +11. principal 在没有相应 target authority 时,不能发现、读取或修改该 target 的标签。 +12. publication 到另一个 Scope 时不创建目标 assignment,也不能隐式暴露 raw source tag。 +13. raw tag value 不进入 telemetry,Dashboard 将恶意值作为 inert text 渲染。 +14. 现有未过滤 API 行为和精确历史 content response 保持不变。 +15. schema、repository、HTTP contract、generated Client 与受支持 backend test 通过仓库标准命令。 + +# Drawbacks + +一张多态 assignment table 无法使用一个普通 conditional foreign key 同时校验 Artifact 和 Memory entry target。 +Repository 必须在 transaction 内强制 Memory-entry 存在。universal catalog-item registry 可以提供一个外键,但会在没有 +其他功能需要它之前增加另一层持久身份和另一张表。 + +无法重建过去某一时刻的 current logical tag。精确 Artifact content 仍可复现,但 tag set 只是 current catalog state。 +需要历史 taxonomy audit 的 deployment 需要额外 event 或 history 设计。 + +两个 query index 会增加写入和存储成本。每个 target 有界的标签数量让成本可预测,而且标签写入频率预计远低于读取。 + +每个受支持 Memory search backend 都必须实现 tag pre-filtering。虽然这比过滤最终 hit 工作量更大,但它是保证 top-k +正确性的必要条件。 + +# Rationale and alternatives + +## 把标签保存在 Artifact content 中 + +该方案可以获得不可变历史标签,但会把一次分类修改变成 content Revision,并在可复用知识没有变化时改变 content digest、 +lineage expectation、CAS behavior 和 derived index。因此不用于 user-managed catalog tag。 + +## 给 `pc_artifact_heads` 增加 JSON 标签列 + +对于整个 Artifact,这只是一个较小的 schema change,但它不能区分标准 one-Memory-per-Scope 模型中的单条 entry。 +此外,SQLite 与 MySQL-compatible backend 对可移植的 indexed `all`/`any` query 支持不同。因此拒绝该方案。 + +## 分别给 Artifact 和 Memory entry 增加一张表 + +该方案具有直接 foreign key 和简单的 family-local join,但会分裂 cross-family query,并重复标签规范化、mutation、 +pagination 与 API behavior。单张多态表保留一个 assignment contract,同时接受一个显式 application-level Memory-entry +invariant。 + +## 增加 universal catalog-item registry + +registry 可以给每个 nested 与 top-level resource 一个统一 ID,并让标签只引用一个 parent table,但它需要为所有现有资源 +新增 lifecycle、migration、ownership、deletion 与 synchronization 语义。第一阶段标签能力不足以证明该抽象的必要性。 + +## 把每条 Memory entry 重构为 Artifact + +该方案会让物理标签 target 统一,却会替换现有 Memory Manifest、atomic collection Revision、entry-version、citation、 +flush 和 index 模型。相对客户需求改动过大,因此拒绝。 + +## 把标签实现为任意 metadata key/value + +当前需求是集合成员关系与精确过滤。通用 nested metadata object 需要额外的 type、operator、indexing、conflict 与 +authorization 语义。本 RFC 不需要这些能力;后续 key/value metadata 功能不能静默重新解释标签。 + +## 不做任何改变 + +用户只能继续把分类编码进 content、维护外部表格或依赖有歧义的文本查询。这些方案都无法提供一致、Scope-aware、 +cross-family 的检索 contract。 + +# Prior art + +PowerContext 已经将不可变 Artifact Revision 与可变 current head、可重建 search projection 分开。RFC 0014 定义 Memory +entry identity 与 exact citation;RFC 0019 定义标准 profile 中每个 Scope 一个 current Memory Artifact;RFC 0048 将 +Handoff 定义为 self-contained Artifact lifecycle;RFC 0051 将 Experience 与 managed Skill 定义为独立 Artifact +Family。本 RFC 把 logical identity 与 immutable content 的同一分离原则应用到 user-managed classification。 + +RFC 1396 将 resource authorization 与 Artifact content 分离,并强调 Scope identity 不是 authorization。标签遵守该边界: +它可以帮助用户找到资源,但永远不能决定用户是否可以访问资源。 + +RFC 1437 建立 Scope-owned Artifact URI tree、mutable current representation 的 opaque HTTP validator,以及绑定 caller +与 query 的 expiring cursor。本 RFC 为该 URI tree 增加 logical-target tag subresource,并扩展 Artifact listing;它不改变 +exact Revision response,也不把 `tag_digest` 当成 Artifact ETag。 + +本 RFC 不以任何外部系统为规范依据。常见代码仓库和 issue tracker 表明,可变 label 可以组织 immutable 或 versioned +content;但 PowerContext 的 Memory-entry containment 与 exact citation model 需要本文定义的 target contract。 + +# Unresolved questions + +没有未决问题阻塞第一阶段 contract 的接受。 + +实现 PR 仍需确认 bounded `all` filter 的 backend-specific query plan,并选择符合仓库命名长度限制的精确 index name。 +这些属于实现验证细节,不能改变公开 normalization、matching、ordering 或 pre-top-k filtering 语义。 + +以下问题刻意排除在本 RFC 之外: + +- 组织是否需要 managed tag definition、颜色、描述、别名或 rename operation; +- tag mutation 是否需要可独立委派的 authorization action; +- 某些 publication workflow 是否应显式提供复制所选标签的选项; +- 自动分类是否可以安全地提议、但不能静默分配标签; +- 是否需要完整 historical tag-assignment ledger。 + +# Future possibilities + +后续 RFC 可以增加 Scope-local tag catalog,提供描述、展示颜色、别名、usage count、受控 rename 或委派 taxonomy +management。该 catalog 用于描述标签;`pc_artifact_tags` 仍然是 assignment relation。 + +另一个扩展可以允许用户保存由 tag、family、lifecycle state 和 text query 组成的 named search view。saved view 必须 +保持为 query,不能成为 authorization policy。 + +模型辅助分类可以通过 reviewed Candidate-like flow 提议标签。模型不能静默分配标签,在授权用户接受前,提议标签保持 +untrusted。 + +如果多个 nested resource type 都需要 tag、access control、favorite、comment 和其他 catalog metadata,项目届时可以 +考虑 universal logical-resource registry。该决策应基于多个经过验证的用例,而不是由本 RFC 提前投机引入。 From 84c05db6a50ef2d9c18e46e5f915a2e3f4edeecc Mon Sep 17 00:00:00 2001 From: Teingi Date: Sat, 5 Sep 2026 01:59:55 +0800 Subject: [PATCH 2/3] docs(rfc): number artifact tags RFC --- docs/en/rfcs/{0000_artifact_tags.md => 1467_artifact_tags.md} | 2 +- docs/zh/rfcs/{0000_artifact_tags.md => 1467_artifact_tags.md} | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) rename docs/en/rfcs/{0000_artifact_tags.md => 1467_artifact_tags.md} (99%) rename docs/zh/rfcs/{0000_artifact_tags.md => 1467_artifact_tags.md} (99%) diff --git a/docs/en/rfcs/0000_artifact_tags.md b/docs/en/rfcs/1467_artifact_tags.md similarity index 99% rename from docs/en/rfcs/0000_artifact_tags.md rename to docs/en/rfcs/1467_artifact_tags.md index bc952ef38..1e9d5965a 100644 --- a/docs/en/rfcs/0000_artifact_tags.md +++ b/docs/en/rfcs/1467_artifact_tags.md @@ -1,6 +1,6 @@ - Proposal Name: `unified_artifact_tags` - Start Date: 2026-09-05 -- RFC PR: [oceanbase/powercontext#0000](https://github.com/oceanbase/powercontext/pull/0000) +- RFC PR: [oceanbase/powercontext#1467](https://github.com/oceanbase/powercontext/pull/1467) - Related Discussion: [oceanbase/powercontext#1466](https://github.com/oceanbase/powercontext/issues/1466) - Related RFCs: [RFC 0014](0014_memory_layer_design.md), [RFC 0019](0019_local_source_memory_runtime.md), [RFC 0048](0048_handoff_artifact.md), [RFC 0051](0051_experience_skill_artifact_families.md), diff --git a/docs/zh/rfcs/0000_artifact_tags.md b/docs/zh/rfcs/1467_artifact_tags.md similarity index 99% rename from docs/zh/rfcs/0000_artifact_tags.md rename to docs/zh/rfcs/1467_artifact_tags.md index bedbfc4ed..3ecf9b41e 100644 --- a/docs/zh/rfcs/0000_artifact_tags.md +++ b/docs/zh/rfcs/1467_artifact_tags.md @@ -1,6 +1,6 @@ - Proposal Name: `unified_artifact_tags` - Start Date: 2026-09-05 -- RFC PR: [oceanbase/powercontext#0000](https://github.com/oceanbase/powercontext/pull/0000) +- RFC PR: [oceanbase/powercontext#1467](https://github.com/oceanbase/powercontext/pull/1467) - Related Discussion: [oceanbase/powercontext#1466](https://github.com/oceanbase/powercontext/issues/1466) - Related RFCs: [RFC 0014](0014_memory_layer_design.md)、[RFC 0019](0019_local_source_memory_runtime.md)、 [RFC 0048](0048_handoff_artifact.md)、[RFC 0051](0051_experience_skill_artifact_families.md)、 From b962344cbacbba50c0b4c76a046e61d693be8796 Mon Sep 17 00:00:00 2001 From: Teingi Date: Sat, 5 Sep 2026 10:52:23 +0800 Subject: [PATCH 3/3] docs(rfc): validate memory tag targets against manifests --- docs/en/rfcs/1467_artifact_tags.md | 53 +++++++++++++++++++++++------- docs/zh/rfcs/1467_artifact_tags.md | 45 ++++++++++++++++++------- 2 files changed, 75 insertions(+), 23 deletions(-) diff --git a/docs/en/rfcs/1467_artifact_tags.md b/docs/en/rfcs/1467_artifact_tags.md index 1e9d5965a..d3d8a926f 100644 --- a/docs/en/rfcs/1467_artifact_tags.md +++ b/docs/en/rfcs/1467_artifact_tags.md @@ -291,10 +291,16 @@ The primary key supports loading one target's tags. The first secondary index su second supports a cross-family Scope query. Implementations must use binary identity comparison or application-built `tag_key` values rather than depending on database-default case or locale collation. -One table cannot express conditional foreign keys to both `pc_artifact_heads` and `pc_memory_entry_heads`. The owning -Artifact foreign key is enforced for every row. For `memory_entry`, the repository must additionally lock and validate -the current `(scope_id, memory_artifact_id, entry_id)` head in the same transaction before changing assignments. The -transaction rejects a missing entry. This is an explicit application invariant, not a best-effort cleanup rule. +The owning Artifact foreign key is enforced for every row. For `memory_entry`, the repository must lock the owning +`(scope_id, family="memory", artifact_id)` Artifact head, load the exact Revision it points to, and validate `entry_id` +against that Revision's authoritative `MemoryContent.manifest.entries` in the same transaction before changing +assignments. Both `active` and `inactive` manifest entries are valid targets. An entry absent from that manifest is +rejected even if an older Revision or immutable entry-version row contains it. + +`pc_memory_entry_heads` contains only active search projections. Deactivation removes an entry's projection while +retaining its logical identity and content in the authoritative manifest. Tag reads and mutations must not require a +row in that table, and tag assignments must not reference it through a foreign key. Projection cleanup and rebuilding +must leave tag assignments unchanged. Inactive Memory entries and deprecated or retired Artifacts retain their assignments. Tag queries apply the existing visibility and lifecycle selection before returning results. A caller with the target's write authority may reorganize @@ -355,7 +361,8 @@ query( `replace` performs the following steps in one transaction: 1. Resolve and authorize the target without returning hidden existence details. -2. Lock the owning Artifact head. For a Memory entry, also lock and validate its current entry head. +2. Lock the owning Artifact head and load the exact Revision it points to. For a Memory entry, validate `entry_id` + against that Revision's manifest, accepting either `active` or `inactive` state. 3. Load the complete current tag set and calculate its digest. 4. Reject a mismatched expected digest as a failed precondition. 5. Validate and normalize the complete replacement set. @@ -369,6 +376,9 @@ compare-and-swap behavior. If the current canonical set equals the requested set, `replace` is an idempotent success and changes no rows or `assigned_at` values. +`get` uses the same authoritative manifest membership rule for Memory entries. An authorized caller can read, +replace, or clear an inactive entry's tags without reactivating it or requiring a search projection. + ## HTTP contract The first contract adds five operations below the Scope resource tree established by RFC 1437: @@ -452,6 +462,12 @@ by `(family, target_type, artifact_id, target_id)` using UTF-8 byte order. The o tag keys, match mode, selected families, target types, lifecycle selection, caller, expiration, and last ordering key. An invalid or filter-mismatched cursor returns `400 Bad Request`; an expired cursor returns `410 Gone`. +For each Memory entry target matched by the tag assignments, resolve the owning Artifact head once and load that +exact Revision's manifest. Use the manifest entry's state for lifecycle selection and its `entry_version_id`, together +with that same Memory Revision, to build the citation. With `include_inactive=true`, eligible inactive entries remain +discoverable without a `pc_memory_entry_heads` row. Do not use an inner join to active projections to determine their +existence or citation. Apply manifest membership, lifecycle selection, and authorization before the page limit. + Tag mutations between pages can change membership. Pagination guarantees deterministic keyset traversal for each query but does not claim a database snapshot across requests. @@ -470,6 +486,10 @@ at least one tag. `tag_match` without `tag` is invalid. Artifact-head listing on Memory-entry list and search only match `memory_entry` targets. Artifact-list cursors additionally bind the normalized tags and match mode; the RFC 1437 invalid, mismatched, and expired cursor statuses remain unchanged. +Memory-entry listing with `include_inactive=true` uses the same manifest-based membership, state, and citation rules +as the dedicated tag query. Memory search retains its existing active-entry eligibility; inactive catalog discovery +does not make an entry searchable. + Existing Artifact-list, Memory-entry-list, and Memory-search item schemas do not gain a `tags` field; their tag filter changes eligibility only. The dedicated tag query includes current tags, and callers can GET the logical target's tag subresource when current catalog metadata is required. Exact historical Artifact Revision and Memory entry-version @@ -574,13 +594,21 @@ The RFC is implemented only when all of the following observable scenarios pass: 14. Existing unfiltered API behavior and exact historical content responses remain unchanged. 15. Schema, repository, HTTP contract, generated-client, and supported backend tests pass through repository-standard commands. +16. After a tagged Memory entry is deactivated and its search projection disappears, an authorized caller can still + read and replace its tags. A matching tag query and Memory-entry list with `include_inactive=true` return it with + a citation resolved from the current manifest; their default requests and Memory search exclude it. Rebuilding + projections preserves these behaviors and assignments. Clearing the tags then returns an empty tag set, removes + it from matching tag queries, and leaves its inactive state, Memory Revision, and entry version unchanged. +17. A Memory-entry tag target absent from the current manifest is rejected even when its immutable entry version + remains stored. An authorized GET of such a target returns `404`, replacement creates no assignments, and tag + queries omit it. # Drawbacks -A polymorphic assignment table cannot use one ordinary conditional foreign key to validate both Artifact and Memory -entry targets. The repository must enforce Memory-entry existence transactionally. A universal catalog-item registry -would provide a single foreign key, but it would add another persistent identity layer and another table before any -other feature needs one. +A relational foreign key validates the owning Artifact, but it cannot enforce logical Memory-entry membership in +that Artifact's current manifest. The repository must enforce that membership transactionally. A universal +catalog-item registry would provide a single foreign key, but it would add another persistent identity layer and +another table before any other feature needs one. Current logical tags are not reconstructible at an earlier time. Exact Artifact content remains reproducible, but the tag set is only current catalog state. Deployments requiring historical taxonomy audit need an additional event or @@ -608,9 +636,10 @@ backends. It is rejected. ## Add one table for Artifacts and another for Memory entries -This provides direct foreign keys and simple family-local joins. It fragments cross-family queries and duplicates tag -normalization, mutation, pagination, and API behavior. The single polymorphic table keeps one assignment contract and -accepts one explicit application-level Memory-entry invariant. +Separate assignment tables simplify family-local joins but fragment cross-family queries and duplicate tag +normalization, mutation, pagination, and API behavior. A separate Memory-entry tag table would still need current +manifest validation unless a durable logical-entry registry were also added. The single polymorphic table keeps one +assignment contract and the same explicit application-level Memory-entry invariant. ## Add a universal catalog-item registry diff --git a/docs/zh/rfcs/1467_artifact_tags.md b/docs/zh/rfcs/1467_artifact_tags.md index 3ecf9b41e..9a973f08f 100644 --- a/docs/zh/rfcs/1467_artifact_tags.md +++ b/docs/zh/rfcs/1467_artifact_tags.md @@ -277,10 +277,14 @@ pc_artifact_tags primary key 支持加载一个 target 的标签;第一个 secondary index 支持 family-specific 过滤,第二个支持 Scope 内跨 Family 查询。实现必须使用 binary identity comparison 或应用构造的 `tag_key`,不能依赖数据库默认大小写或 locale collation。 -一张表无法用普通 conditional foreign key 同时校验 `pc_artifact_heads` 和 `pc_memory_entry_heads`。每一行都强制所属 -Artifact foreign key。对于 `memory_entry`,repository 还必须在同一个 transaction 中锁定并校验当前 -`(scope_id, memory_artifact_id, entry_id)` head,再修改分配关系;缺少 entry 时 transaction 拒绝请求。这是明确的 -application invariant,不是 best-effort cleanup rule。 +每一行都强制所属 Artifact foreign key。对于 `memory_entry`,repository 必须在同一个 transaction 中锁定所属 +`(scope_id, family="memory", artifact_id)` Artifact head,加载其指向的精确 Revision,并根据该 Revision 的权威 +`MemoryContent.manifest.entries` 校验 `entry_id`,再修改分配关系。manifest 中 `active` 和 `inactive` entry 都是有效 +target。当前 manifest 中不存在的 entry 必须被拒绝,即使旧 Revision 或不可变 entry-version row 中仍保留该 entry。 + +`pc_memory_entry_heads` 只包含 active search projection。停用 entry 会删除它的 projection,但其逻辑身份与内容仍保留在 +权威 manifest 中。标签读取与变更不能要求该表中存在对应 row,标签分配也不能通过 foreign key 引用该表。projection +清理与重建必须保持标签分配不变。 inactive Memory entry 与 deprecated 或 retired Artifact 保留其分配关系。标签查询先应用现有 visibility 和 lifecycle selection,再返回结果。拥有 target write authority 的调用方可以重新组织标签,而不会重新激活或修订 content。 @@ -339,7 +343,8 @@ query( `replace` 在一个 transaction 内执行以下步骤: 1. 解析并授权 target,同时不返回隐藏的存在性细节。 -2. 锁定所属 Artifact head;对于 Memory entry,还要锁定并校验其 current entry head。 +2. 锁定所属 Artifact head,加载其指向的精确 Revision;对于 Memory entry,根据该 Revision 的 manifest 校验 + `entry_id`,接受 `active` 或 `inactive` 状态。 3. 加载完整 current tag set 并计算 digest。 4. expected digest 不匹配时,按 precondition failure 拒绝请求。 5. 校验并规范化完整 replacement set。 @@ -351,6 +356,9 @@ process-local lock,因为二者都不能提供所需的分布式 compare-and-s 如果当前 canonical set 与请求集合相同,`replace` 幂等成功,并且不修改任何 row 或 `assigned_at` 值。 +`get` 对 Memory entry 使用相同的权威 manifest 成员校验规则。授权调用方可以读取、替换或清空 inactive entry 的标签, +无需重新激活 entry,也不要求存在 search projection。 + ## HTTP contract 第一阶段 contract 在 RFC 1437 建立的 Scope resource tree 下新增五个 operation: @@ -434,6 +442,12 @@ Artifact target 的 `current` 使用 `ArtifactReference`;Memory entry target key、match mode、所选 family、target type、lifecycle selection、caller、expiration 和最后一个 ordering key。非法或 filter 不匹配的 cursor 返回 `400 Bad Request`;过期 cursor 返回 `410 Gone`。 +对于每个标签分配命中的 Memory entry target,只解析一次所属 Artifact head,并加载该精确 Revision 的 manifest。使用 +manifest entry 的 state 进行 lifecycle selection,再用它的 `entry_version_id` 和同一个 Memory Revision 构造 citation。 +当 `include_inactive=true` 时,符合条件的 inactive entry 即使没有 `pc_memory_entry_heads` row,也仍可被发现。不能通过 +与 active projection 的 inner join 判断其存在性或解析 citation。manifest 成员校验、lifecycle selection 和 authorization +必须在 page limit 之前应用。 + page 之间发生的标签变更可能改变成员关系。pagination 为每次查询提供确定性 keyset traversal,但不承诺跨请求数据库 snapshot。 @@ -450,6 +464,9 @@ parameter 或 field 默认缺失,因此保留现有行为。filter 存在时 非法请求。Artifact-head listing 只匹配 `artifact` target;Memory-entry list 和 search 只匹配 `memory_entry` target。 Artifact-list cursor 还要绑定 normalized tag 与 match mode;RFC 1437 的非法、不匹配和过期 cursor status 保持不变。 +Memory-entry listing 在 `include_inactive=true` 时,使用与专用 tag query 相同的 manifest 成员校验、state 与 citation +规则。Memory search 保持现有的 active-entry eligibility;inactive 条目的 catalog discovery 不会使它进入搜索结果。 + 现有 Artifact-list、Memory-entry-list 和 Memory-search item schema 不新增 `tags` 字段;tag filter 只改变 eligibility。 专用 tag query 包含 current tags;需要 current catalog metadata 时,调用方可以 GET logical target 的 tag subresource。 精确历史 Artifact Revision 和 Memory entry-version response 不新增 `tags` 字段。 @@ -546,12 +563,18 @@ table;需要完整标签变更账本的 deployment 必须禁用该能力,或 13. raw tag value 不进入 telemetry,Dashboard 将恶意值作为 inert text 渲染。 14. 现有未过滤 API 行为和精确历史 content response 保持不变。 15. schema、repository、HTTP contract、generated Client 与受支持 backend test 通过仓库标准命令。 +16. 带标签的 Memory entry 停用且其 search projection 消失后,授权调用方仍能读取和替换它的标签。匹配的 tag query 与 + Memory-entry list 在 `include_inactive=true` 时返回该 entry,citation 从 current manifest 解析;默认请求和 + Memory search 不返回它。重建 projection 后,这些行为与标签分配保持不变。随后清空标签会返回空 tag set,使该 entry + 不再命中相应 tag query,同时保持其 inactive 状态、Memory Revision 和 entry version 不变。 +17. 当前 manifest 中不存在的 Memory-entry tag target 会被拒绝,即使其不可变 entry version 仍保留在存储中。授权 GET + 返回 `404`,replacement 不创建 assignment,tag query 不返回该 target。 # Drawbacks -一张多态 assignment table 无法使用一个普通 conditional foreign key 同时校验 Artifact 和 Memory entry target。 -Repository 必须在 transaction 内强制 Memory-entry 存在。universal catalog-item registry 可以提供一个外键,但会在没有 -其他功能需要它之前增加另一层持久身份和另一张表。 +关系型 foreign key 可以校验所属 Artifact,但无法强制 logical Memory entry 属于该 Artifact 的 current manifest。 +Repository 必须在 transaction 内校验该成员关系。universal catalog-item registry 可以提供一个外键,但会在没有其他功能 +需要它之前增加另一层持久身份和另一张表。 无法重建过去某一时刻的 current logical tag。精确 Artifact content 仍可复现,但 tag set 只是 current catalog state。 需要历史 taxonomy audit 的 deployment 需要额外 event 或 history 设计。 @@ -575,9 +598,9 @@ lineage expectation、CAS behavior 和 derived index。因此不用于 user-mana ## 分别给 Artifact 和 Memory entry 增加一张表 -该方案具有直接 foreign key 和简单的 family-local join,但会分裂 cross-family query,并重复标签规范化、mutation、 -pagination 与 API behavior。单张多态表保留一个 assignment contract,同时接受一个显式 application-level Memory-entry -invariant。 +分开的 assignment table 可以简化 family-local join,但会分裂 cross-family query,并重复标签规范化、mutation、 +pagination 与 API behavior。除非再增加持久的 logical-entry registry,否则独立的 Memory-entry tag table 仍需要校验 +current manifest。单张多态表保留一个 assignment contract,并承担相同的显式 application-level Memory-entry invariant。 ## 增加 universal catalog-item registry