What was run, in order, with the numbers behind each result. Every entry comes from the
experiment ledger in the repository; final-test, validation and development evidence are labelled separately.
24 runs2 final-test runs24 Aug 2026 – 27 Aug 2026
Final testfrozen config, run once on held-out testValidationthree seeds, test disabledDevelopmentseed-42 selection or screenProtocolplans, freezes, clean-up
Added a non-trainable, causal user anchor computed only from official pre-behavior training history; all validation impressions have anchor coverage and local-test history is excluded.
EB-NeRD
Full notes
The fixed-anchor/query-age variant confirms validation NDCG@10 0.52557 ± 0.00053 across seeds 42/43/44, above the earlier selected FWPRec result 0.51696 ± 0.00206.
Replacing query-age features with cache-invariant inter-event gaps gives 0.51903 ± 0.00424; this is a separate, less stable quality result.
Exact aligned horizons 10/20/50 give seed-42 NDCG@10 0.51348/0.51293/0.51519. Horizon 50 has 100% top-1 cached/reconstructed agreement over 200 unchanged real pairs with score MAE 9.74e-8.
Conclusion: fixed historical anchoring is a robust quality improvement, but exact hard-reset sessions pay a ranking cost. The strongest quality and strongest exact-persistence settings remain separate paper rows. No test split was evaluated.
Full protocol and outputs: evaluation/EBNERD_PERSISTENT_FWPREC.md.
Added an optional causal user anchor built from positive frozen-Qwen items in each user's earliest training example; validation and test histories are not used.
MINDAmazon
Full notes
Under a matched seed-42 anchor-by-memory factorial, fixed-anchor FWP improves MIND validation NDCG@10 from 0.36542 to 0.38907.
On Amazon it decreases full FWP NDCG@10 from 0.26700 to 0.26309, retained as a negative transfer result. Fixed anchor alone nevertheless improves over reconstructed anchor alone, and FWP improves strongly over the fixed-anchor control.
The paper must distinguish an exogenous interest/query anchor from a historical proxy. Stability is an architectural assumption; a particular historical mean is not guaranteed to be the best anchor on every domain.
No test split was evaluated. Full results and reproduction commands are in evaluation/FIXED_ANCHOR_CROSS_DATASET_SCREEN.md.
Added a clean-room reader for official behavior, history, and article Parquet schemas with chronological, causal state construction.
EB-NeRD
Full notes
Kept clicked-only history as the default; exposed non-click feedback is an optional separately labeled protocol.
Added a shared frozen-Qwen representation path and a six-model validation-only smoke suite. No hidden test or model-selection result has been accessed.
Downloaded the official small archive supplied by the user and verified SHA-256 84fc6fc902b8e8dad8383835c90185d54e0916060586d428ced391da4a92bd4f.
The model-free audit found 232,887 chronological training impressions, 244,647 labeled-validation impressions, 20,738 articles, complete logged candidate/click consistency, and no fabricated events or labels. Full audit: evaluation/EBNERD_DATASET_AUDIT.md.
Frozen article encoding and the validation-only interface smoke are the next stage; no hidden test or EB-NeRD model tuning has been run.
Encoded all 20,738 articles once with the pinned 64-dimensional Qwen setup; verified unique IDs, finite unit vectors, source provenance, and archive SHA-256 d3da47dcb1f6925fb43b23bea751e7decef6bd446c51aaf03f05b6679bbee43e.
EB-NeRD
Full notes
All six smoke models completed on the same chronological 45,000/5,000 development split with clicks-only histories and evaluate_test: false.
PositiveEMA and SASRecProfile reached diagnostic NDCG@10 of 0.4919 and 0.4918; FWPRec reached 0.4825. These one-epoch, one-seed prefix results are not paper comparisons.
FWPRec's positive memory was active and negative memory remained zero as expected under clicks-only updates. Details: evaluation/EBNERD_QWEN_SMOKE.md.
Prespecified next step: full-training-split one-seed convergence development, followed by bounded tuning and a frozen three-seed confirmation. The official labeled validation split remains uninspected.
Completed nine shared-Qwen comparators on 209,598 training and 23,289 chronological development-validation impressions; evaluate_test: false.
EB-NeRD
Full notes
FWPRec NDCG@10 was 0.5123, close to HSTU at 0.5135 and SASRecProfile at 0.5164, and above GRU4Rec, LRURec, DuoRec, PositiveEMA, and Popularity in this single-seed development run.
FWPRec alone achieved its best selected metric at the 10-epoch boundary, so it receives an unchanged 15-epoch extension. This is a convergence check, not model-specific tuning based on relative rank.
Full table and limitations: evaluation/EBNERD_QWEN_DEVELOPMENT.md.
The unchanged 15-epoch FWPRec extension peaked at epoch 12 with NDCG@10 0.51525, closing the convergence budget and reducing the seed-42 gap to SASRecProfile to 0.00118. Next: symmetric bounded learning-rate screen.
Completed the symmetric {0.0005, 0.001, 0.002} learning-rate screen for every learned comparator. FWPRec selected 0.002 at NDCG@10 0.51655; no EB-NeRD-specific FWP memory-architecture search will be performed. Frozen confirmation adds seeds 43/44 with test evaluation disabled. Details: evaluation/EBNERD_QWEN_TUNING.md.
Froze per-model learning rates from the common three-value grid and added seeds 43/44; seed 42 was reused from selection. No test evaluation occurred.
EB-NeRD
Full notes
Across seeds 42/43/44, FWPRec achieved AUC 0.60029 ± 0.00413 and NDCG@10 0.51696 ± 0.00206, compared with SASRecProfile at 0.51413 ± 0.00355 and HSTU at 0.51260 ± 0.00080.
FWPRec exceeded HSTU, GRU4Rec, and LRURec on every seed and SASRecProfile on two of three. This is descriptive validation evidence, not a universal superiority or formal significance claim.
Next EB-NeRD blocks: real cohort/adaptation analysis and trained-checkpoint cost measurements. Full controlled table: evaluation/EBNERD_QWEN_CONFIRMATION.md.
Completed clean-room NRMS and LSTUR under the same real chronological split, logged candidates, clicked-only histories, and validation evaluator; the official labeled validation/local-test split remained untouched.
EB-NeRD
Full notes
Used a common learning-rate grid {0.0005, 0.001, 0.002} on seed 42 and selected 0.002 for both models, then added frozen seeds 43 and 44.
Three-seed validation NDCG@10 was 0.51011 ± 0.00504 for NRMS and 0.49683 ± 0.00264 for LSTUR. Full metrics and provenance: evaluation/EBNERD_NATIVE_NEWS_BASELINES.md.
This is a separately labeled architecture-native comparison. Because these baselines learn title encoders from scratch while FWPRec uses frozen Qwen vectors, their difference does not isolate the fast-weight mechanism and is not a universal superiority result.
Next EB-NeRD blocks remain real-data adaptation analysis and matched trained-checkpoint cost measurement, with no synthetic shifts.
Re-evaluated 25 exact selected shared-Qwen checkpoints without retraining. Ordinary AUC, MRR, NDCG@5, and NDCG@10 reproduced within 1e-6 for every checkpoint, confirming that annotations did not alter predictions or labels.
EB-NeRD
Full notes
Used the previously established causal five-click-window, JSD >= 0.3 detector and unchanged recorded clicks. Validation contains 779 complete H=3 events.
FWPRec achieved H=3 AdaptNDCG@10 0.51066 ± 0.00405, compared with SASRecProfile at 0.50904 ± 0.00550, HSTU at 0.50807 ± 0.00078, and GRU4Rec at 0.50340 ± 0.00220.
FWPRec leads the mean but not every seed against SASRecProfile or HSTU. Report competitive short-horizon adaptation, not formal or universal superiority.
Full protocol, table, and checkpoint hashes: evaluation/EBNERD_REAL_ADAPTATION.md.
Replayed 200 unchanged consecutive single-click transitions through seven exact selected checkpoints on an RTX 4080 SUPER; no training or test evaluation occurred.
EB-NeRD
Full notes
Dense FWP uses 33,024 persistent tensor bytes/user and 104,704 counted update MACs, versus cached GRU at 256 bytes and 24,576 MACs.
FWP cached update+score p50 was 0.7898 ms, modestly below SAS/HSTU full reconstruction but 7.14x cached GRU latency. Do not claim broad efficiency leadership over compact recurrent state.
One-step top-1 agreement with fresh offline reconstruction was 89.5% for FWP and 98.5% for GRU. FWP drift follows its moving slow anchor, 20-write window, and re-aged time features; a persistent formulation or refresh policy must be trained/evaluated before a deployment-equivalence claim.
Full state, MAC, latency, cohort, hardware, and hash record: evaluation/EBNERD_CHECKPOINT_EFFICIENCY.md.
Defined the target as gradient-free bounded-state online adaptation, not universal speed or storage leadership over sequential models.
Full notes
Prespecified operation-separated p50/p95/p99 latency, batch/candidate/history scaling, trained-checkpoint cached replay, TTL state residency, cold rebuild, and a quality-cost Pareto analysis.
Retained cached GRU as the strongest compact systems comparator and required cached FWP to preserve both overall and negative-rich quality relative to exact reconstruction before making the deployment claim.
Full design and decision gates: evaluation/INDUSTRY_SERVING_EXPERIMENT_PLAN.md.
The committed preflight passed all five hashes, suite counts, output-absence, and clean-worktree checks before test access.
Amazon passed · KuaiRand failedKuaiRandAmazon
Full notes
All 19 KuaiRand and 25 Amazon jobs completed in their original suite runs; no failed or selectively rerun seed occurred.
KuaiRand primary joint claim failed: FWPRec versus GRU4Rec overall relative effect -1.720% (95% CI [-2.481%, -.907%]) and complete-event H=3 adaptation difference -.01089 (95% CI [-.02477, .00347]).
Amazon primary joint claim passed: FWPRec versus LRURec overall relative effect +8.719% (95% CI [+7.115%, +10.418%]) and negative-rich difference +.02727 (95% CI [.01814, .03622]).
Compact dense passed secondary overall non-inferiority on both datasets; rank 16 passed on KuaiRand and failed on Amazon. Full provenance and tables: evaluation/CROSS_DATASET_FINAL_TEST_RESULTS.md.
Froze 19 KuaiRand and 25 Amazon final-test jobs using only validation-selected models and hyperparameters. Test evaluation and prediction archiving are enabled only in these named suites.
KuaiRandAmazon
Full notes
Precommitted primary and secondary endpoint requirements and a generalized paired hierarchical-bootstrap analyzer before neural-model test access.
KuaiRand's held-out test split remains untouched. No Amazon neural model has been evaluated on test; an earlier parameter-free Popularity protocol audit did inspect Amazon test performance and is disclosed in the frozen plan.
Added a clean-worktree/hash preflight. Final execution remains blocked until this protocol state is committed without generated result directories.
Added a deterministic artifact generator over frozen MIND prediction archives and matched Amazon five-trial efficiency measurements.
MINDAmazon
Full notes
Generated support-aware adaptation curves, the Amazon accuracy/state/latency figure, CSV tables, and a booktabs LaTeX table in results/paper_artifacts/mind_final/.
Recorded SHA-256 hashes for all source and generated files in manifest.json; no training or post-test model selection was performed.
Added and executed experiments/notebooks/paper_figures_and_tables.ipynb for later visual and table-layout changes. Full semantics: evaluation/PAPER_ARTIFACTS.md.
Verified a clean committed worktree and both frozen YAML hashes before the first test access; no prior final-test directory existed.
Primary claim failedMIND
Full notes
Completed all 37 jobs on the first attempt without overrides, selective execution, or reruns. Summary: results/batch_summaries/mind_final_test_frozen/20260825T151230.534379Z.
Ran the locked 10,000-replicate analysis on 50,000 users and 3,472 complete H=3 events with exact archive alignment.
Standalone FWPRec failed overall non-inferiority (-2.97%, 95% CI [-3.49%, -2.40%]) and adaptation superiority (-.01353, 95% CI [-.01772, -.00933]); the joint primary claim failed.
Secondary SASRecFWP passed overall non-inferiority (+.20%, 95% CI [-.19%, +.62%]) but not adaptation superiority (+.00108, 95% CI [-.00159, +.00385]).
No post-test model, margin, endpoint, or analysis change was made. Full table and interpretation: evaluation/MIND_FINAL_TEST_RESULTS.md.
Froze 37 runs: one deterministic Popularity row and twelve neural methods over seeds 42/43/44 with dataset/candidate seed 42.
MIND
Full notes
Fixed standalone FWPRec versus SASRecProfile as the sole confirmatory comparison; retained the pre-existing 1% overall non-inferiority margin and complete-event H=3 adaptation superiority requirement.
Added safe per-impression NPZ archives and a deterministic 10,000-replicate hierarchical paired bootstrap over users/events and model seeds.
Dry-run expansion passed; every final job evaluates test once and archives predictions, while controlled/native representation assignments remain separate. No final-test output directory was created.
Frozen suite SHA-256: 51bc852b25a7e910395ad9c4f9475ec942a2231202b4d8c4b39adcde4f19e69a.
Confirmed the Amazon-selected shared write trunk on MIND and KuaiRand with fixed dataset/candidate seed 42, model seeds 42/43/44, and no test access.
MINDKuaiRandAmazon
Full notes
On MIND, compact FWPRec is 0.48% lower overall and 2.84% lower on H=3 adaptation than independent generators. Because adaptation is central and only 25 complete events support it, rank 16 was not stacked on MIND.
On KuaiRand, compact FWPRec is 0.30% lower overall and 0.06% lower on adaptation than the independent model; the first-step mean is higher.
KuaiRand compact rank 16 retains 99.86% of compact-dense overall nDCG@10 and 99.81% of its AdaptNDCG while reducing dual-memory state by 49.6%.
The supported result is a transferable generator reduction and a credible accuracy/state Pareto point, not universal adaptation equivalence or a speed improvement. Exact artifacts are in evaluation/CROSS_DATASET_COMPACT_FWPREC.md.
Completed matched learning-rate screens for HSTU, LRURec, and FWPRec under the frozen-Qwen 100-candidate protocol, followed by one-factor FWPRec recent context and negative-strength screens.
Amazon
Full notes
Selected learning rates .001, .002, and .001 for HSTU, LRURec, and FWPRec, respectively, using overall validation nDCG@10 only.
Provisional FWPRec selection is recent_history_length=20 and alpha_negative=.5; its nDCG@10 is .26916 versus HSTU's .27502.
FWPRec leads HSTU on negative-rich nDCG@10 (.25595 versus .23996), but not on overall, adaptation, first-step, avoidance, or long-history endpoints.
The .5 versus .3 alpha difference is only 0.15% relative and must be confirmed over seeds 42/43/44 before any test evaluation.
All nine new neural runs completed successfully on CUDA; test evaluation remained disabled.
The retained corrected confirmation pins dataset/candidate seed 42 for every model seed.
Interpretation: NRMS supplies a strong architecture-native overall baseline; standalone FWPRec retains a slightly higher confirmed adaptation mean (0.41991) but is not representation-controlled against NRMS.
Added explicit dual, positive_only, and single_signed memory layouts with batch/streaming equivalence tests.
Full notes
Ran five prespecified one-factor variants over seeds 42/43/44 with test evaluation disabled.
Learned value generation produced the largest contribution; learned key addressing produced a smaller consistent contribution.
Evolving write state improved overall nDCG for every seed, from 0.36692 ± 0.00046 to 0.36782 ± 0.00037, but its adaptation gain was not seed-consistent.
Positive-only matched dual memory, so standalone negative-memory benefit is not supported by this experiment.
Both exceed FWPRec's corresponding means; HSTU leads every recorded endpoint on every seed. The earlier strongest-challenger claim based only on GRU4Rec is superseded.
All five new YAML suites passed dry-run expansion, 125 tests passed, and git diff --check passed.
The idea
Research question and current position
The original paper proposes a reciprocal user-matching model in which a user's stated look-for representation is stable and fast-weight memory adjusts it using recent interactions. Because suitable reciprocal user-user datasets were not available, this repository studies the same mechanism under conventional user-to-item candidate ranking and removes retrospective reciprocal matching.
The current defensible position is:
FWPRec is an explicit bounded associative adaptation mechanism. It reuses cacheable item representations and updates user-specific fast state without changing global parameters. It offers a conditional quality/state/computation trade-off, with particularly encouraging evidence under explicit signed feedback. It is not a universal replacement for sequential recommendation, nor is dense FWPRec universally smaller or faster than a cached recurrent model.
Do not reposition the paper as “FWPRec beats sequence models.” Frozen MIND and KuaiRand tests reject that story. The most promising scientific contribution is the separation of a stable slow intent/profile anchor from learned, gradient-free fast associative updates, together with careful mechanism and state/accuracy analysis.
Proposed NeurIPS abstract
Modern sequential recommenders repeatedly encode interaction histories to track evolving user preferences, creating a tension between adaptation quality and online computation. We propose FWPRec, a lightweight alternative that decomposes user state into a stable preference anchor and a rapidly updated associative memory. The anchor may represent an explicit interest, search query, or cached long-term profile, while fast-weight matrices incorporate newly observed positive and negative feedback through gradient-free updates. Item representations and the slow anchor are reusable across requests; adaptation therefore requires neither parameter updates nor repeated sequence encoding. We evaluate FWPRec on real interaction data from MIND, EB-NeRD, KuaiRand, and Amazon Reviews using shared candidate sets, frozen semantic item embeddings, chronological splits, and strong recurrent, attention-based, state-space, and non-sequential baselines. FWPRec provides competitive ranking and adaptation performance, including improvements over strong sequential models in several controlled settings, while retaining bounded user state and incremental updates. Ablations show that learned value construction and the separation between a stable anchor and fast preference shifts are important. They also reveal an important limitation: historical averages are imperfect substitutes for explicit intent and do not improve every domain. FWPRec thus offers a scientifically grounded accuracy–adaptation–cost trade-off for recommendation settings where continually reconstructing large sequential user models is undesirable.
This abstract is intentionally conditional. Before submission, its efficiency phrasing should be checked against the dense-state and cached-GRU results in the checkpoint efficiency report.
How the model works
Method snapshot
Let q_base be a stable user interest, search query, or causal slow profile. For event l, slow write networks produce a key, value, and gate from the base state, item embedding, feedback sign/action, and time features. Positive and negative delta-rule memories are updated independently:
The matrices are per-user transient tensors, not learned parameters. Gradients train the slow write networks; an observed interaction updates fast state without an optimizer step. See model.py, memory.py, and the original mathematical discussion.
Implemented factors include:
independent versus compact shared-trunk write networks;
dense versus bounded low-rank factor-ring memory;
dual, positive-only, and single-signed memory layouts;
learned key and value residuals;
fixed, mean-content, attention, learned-user, and hybrid anchors;
optional evolving fast state in each write input;
query-age or invariant inter-event time features; and
batch reconstruction and incremental session-state execution.
The canonical configuration schema is configs/models/fwprec.yaml. Paper experiments use their checked-in resolved YAML settings rather than assuming these defaults.
Evidence hierarchy
Status
Meaning
Examples
Final test
Frozen configuration and analysis executed once
MIND, KuaiRand, Amazon
Validation-confirmed
Three seeds, test disabled
EB-NeRD controlled ranking, mechanism and compression studies
Never promote development or diagnostic values into a final-test table. In particular, do not use synthetic fixtures, retired manipulated preference shifts, learned-ID MIND cold-start smoke runs, or deleted random-weight timing measurements as scientific evidence. The governing real-data policy is REAL_DATA_ADAPTATION_PROTOCOL.md.
Frozen final-test results
Final-test evidence
MIND-small: primary controlled test
All controlled neural models use the same frozen normalized 64-dimensional Qwen article vectors, candidates, split, and dataset seed. Neural results use model seeds 42/43/44. The frozen standalone claim required FWPRec to be within 1% of SASRecProfile overall and superior on complete-event H=3 adaptation. Both requirements failed.
Model
Test NDCG@10
AdaptNDCG@10
GRU4Rec
0.38855 ± 0.00113
0.38055 ± 0.00114
LRURec
0.39287 ± 0.00184
0.38409 ± 0.00087
HSTU
0.39217 ± 0.00020
0.38445 ± 0.00155
SASRecProfile
0.39649 ± 0.00114
0.38921 ± 0.00151
FWPRec
0.38472 ± 0.00158
0.37568 ± 0.00161
SASRecSignedEMA
0.39872 ± 0.00237
0.38754 ± 0.00250
SASRecFWP
0.39729 ± 0.00125
0.39029 ± 0.00179
Standalone FWPRec's overall effect is -2.97% relative to SASRecProfile, with a 95% interval of [-3.49%, -2.40%]. Its adaptation difference is -0.01353 with a 95% interval of [-0.01772, -0.00933]. SASRecFWP is overall non-inferior, but its adaptation interval crosses zero. Signed EMA is at least as compelling as the hybrid FWP control, so hybrid benefits cannot be uniquely attributed to matrix memory.
KuaiRand: negative final result for standalone leadership
FWPRec's prespecified comparison with GRU4Rec failed overall non-inferiority and did not establish adaptation superiority. HSTU is descriptively strongest on every reported endpoint.
Model
Test NDCG@10
AdaptNDCG@10
Negative-rich NDCG@10
GRU4Rec
0.57116 ± 0.00162
0.60072 ± 0.00399
0.45427 ± 0.00204
HSTU
0.61211 ± 0.00579
0.63889 ± 0.00698
0.49985 ± 0.00549
Mamba4Rec
0.59205 ± 0.00631
0.61966 ± 0.00624
0.47829 ± 0.00801
FWPRec, independent dense
0.56133 ± 0.00262
0.58984 ± 0.00735
0.40590 ± 0.00100
Retain this as a valid failed claim. Do not optimize a new KuaiRand variant against the already observed test set.
Amazon Video Games: strongest positive final evidence
Amazon uses explicit signed ratings, frozen Qwen product vectors, chronological splits, 100 candidates, and seeds 42/43/44. FWPRec's joint comparison with LRURec passed both overall non-inferiority and negative-rich superiority.
Model
Test NDCG@10
AdaptNDCG@10
Negative-rich NDCG@10
SignedEMA
0.24390 ± 0.00059
0.29838 ± 0.00024
0.18228 ± 0.00061
LRURec
0.25816 ± 0.00237
0.29986 ± 0.00677
0.22480 ± 0.00175
HSTU
0.26776 ± 0.00254
0.31905 ± 0.00651
0.23192 ± 0.00162
FWPRec, independent dense
0.28067 ± 0.00106
0.32276 ± 0.00356
0.25208 ± 0.00157
FWPRec, compact dense
0.28237 ± 0.00210
0.32804 ± 0.00210
0.25534 ± 0.00125
Against LRURec, the overall relative effect is +8.719% with 95% interval [+7.115%, +10.418%], and the negative-rich difference is +0.02727 with interval [0.01814, 0.03622]. This supports signed-feedback specialization, not universal adaptation superiority.
The complete KuaiRand/Amazon tables and locked analyses are in CROSS_DATASET_FINAL_TEST_RESULTS.md. Do not rerun or revise their frozen hypotheses in response to these outcomes.
Efficiency and persistent-state evidence
The checkpoint-backed EB-NeRD audit uses 200 unchanged consecutive real single-click transitions and a trained seed-42 checkpoint on an RTX 4080 SUPER.
Model
State bytes/user
Counted update MACs
Cached update+score p50
GRU4Rec
256
24,576
0.1106 ms
FWPRec, dense dual
33,024
104,704
0.7898 ms
SASRecProfile, reconstruction
—
—
0.8746 ms
HSTU, reconstruction
—
—
1.0076 ms
Dense FWP is bounded in history length but is neither smaller nor faster than a cached compact GRU at dimension 64. It is modestly faster than full SAS/HSTU reconstruction in this implementation, but those models were not evaluated with optimized KV caches. Wall-clock evidence is secondary.
The original reconstructed FWP checkpoint has only 89.5% one-step top-1 agreement with cached updates because its anchor, recent window, and query-age features change during reconstruction. A fixed anchor with invariant inter-event features and aligned horizon 50 reaches 100% top-1 agreement and score MAE 9.74e-8, but its seed-42 NDCG@10 falls to 0.51519. Thus exact bounded sessions are implemented, but hard resets have a quality cost. The best-quality fixed-anchor model and the exact-persistence model are currently different variants.
MIND validation originally had only 25 complete H=3 events; its apparent adaptation advantage did not survive the much larger frozen test.
Dense FWP state is large compared with cached GRU state and costs O(d^2) per update. “Bounded in history” does not mean “small in dimension.”
Current exact persistent execution uses aligned resets and pays a ranking penalty. Unreset stateful training/evaluation is not yet implemented.
Low-rank state reduces bytes but is not faster in the current unfused code.
EB-NeRD results remain validation-only; do not describe them as held-out test confirmation.
The baseline implementations are clean-room research reproductions, not byte-identical copies of upstream repositories.
MIND/EB-NeRD native news encoders learn word embeddings from scratch and are not directly representation-matched to frozen Qwen models.
Wall-clock measurements are hardware- and implementation-specific and must accompany analytical state and operation counts.
Recommended continuation order
1. Confirm the stable-anchor MIND result
Run seeds 43 and 44 for the already defined MIND fixed-anchor factorial without changing its seed-42-selected settings. Report all four rows, not only fixed anchor + FWP. This is the fastest test of whether the restored original slow/fast decomposition is robust.
Do not treat Amazon as a tuning target for inventing a favorable fixed proxy. Its negative transfer result should remain visible. A new Amazon anchor is scientifically justified only if its construction is motivated independently by the application, such as a prespecified longer pre-period signed profile.
2. Make persistent training match persistent serving
Implement a stateful objective that processes real user events in causal order, keeps a fixed exogenous/slow anchor, uses invariant write-time features, and truncates gradients without resetting fast state. Evaluate both ranking quality and cached/reconstructed semantics. Avoid synthetic transitions and artificial history manipulation.
The main target is to remove the hard-reset quality penalty while retaining exact incremental execution. Tensorize per-user update counters before making batched serving-throughput claims.
3. Re-evaluate the compact and low-rank variants under those semantics
Start from the shared-trunk writer because its parameter reduction is the most stable lightweight result. Compare dense, positive-only where appropriate, and rank-16 state. Report bytes, analytical operations, checkpoint-backed latency, and quality together. A smaller state is useful even if latency does not fall.
4. Obtain evidence with genuine exogenous anchors
The original hypothesis is best tested on a dataset with an explicit search query, stated interest, or session intent available before interaction. Keep that anchor fixed and compare:
anchor only
anchor + EMA
anchor + GRU update
anchor + FWP
strong full sequential model
This would distinguish the FWP mechanism from historical-anchor engineering and reconnect the user-to-item study to the original user-user motivation.
5. Transfer the real checkpoint efficiency audit
Repeat the EB-NeRD cached-state study on Amazon, where FWP has its strongest final quality evidence and signed memory is actually used. Include cached GRU, EMA, dense compact FWP, and rank-16 compact FWP under identical candidates and hardware. Do not claim comparison with an optimized Transformer cache unless one is actually implemented.
6. Update the paper only after the above evidence
Use the existing notebook/artifact pipeline. Preserve separate tables for:
representation-controlled ranking;
signed-feedback and adaptation endpoints;
mechanism ablations;
stable-anchor experiments; and
quality/state/computation trade-offs.
Do not combine native news encoders, frozen-Qwen user-state controls, and different candidate protocols into a single causal ranking.