CacheRegime / Systems research

When should a cache
change policy?

Use a short request prefix to choose a cache-policy family. Then measure whether choosing the right family actually improves cache effectiveness.

59 / 63 correct held-out policy-family predictions
7 traces one trace held out for each evaluation fold
+0.141 pp mean hit-rate gain over the best fixed baseline

A selector is useful only if the cache benefits.

Different workloads reward different eviction policies. I studied whether a short warm-up prefix contains enough information to choose a useful policy family before replaying the rest of a mixed workload.

The experiment combines seven public cache traces with a WebLINX-derived synthetic agent overlay. Each trace is evaluated under three overlay behaviors and three target agent fractions, giving 63 settings at a 64 MiB cache size. The selector profiles the first 1,000 mixed requests. It chooses a rate-guarded representative from WTinyLFU, S3FIFO or GDSF.

The evaluation holds out an entire trace, including all its settings. A small threshold tree is fitted on the other six traces. This asks whether the rule transfers to a different trace within the panel.

1,000-request prefix Workload statistics Fitted threshold tree Guarded cache policy

Measured result

Good family prediction. A modest hit-rate gain.

CacheRegime: selector mean human hit rate 13.09%, best fixed WTinyLFU baseline 12.94%, oracle 13.22%. Held-out family predictions correct in 59 of 63 settings.
Recomputed from 513 archived policy measurements and frozen prefix features. Rates are unweighted means over the 63 settings. Agent traffic is a synthetic overlay; the cache budget is 64 MiB.
Mean human-request hit rate. “Human” denotes the original public-trace component in the mixed workload.
Approach Hit rate How the policy is chosen
Per-setting oracle 13.22% Best measured policy after seeing the outcome
Held-out selector 13.09% Tree trained on the other six traces
Fixed WTinyLFU + rate guard 12.94% Same policy in every setting
Fixed S3FIFO + rate guard 11.77% Same policy in every setting
Fixed GDSF + rate guard 10.67% Same policy in every setting

The selector matches the winning family in 59 of 63 settings, or 93.7%. That accuracy does not translate into a 93.7% cache hit rate. The practical gain over the strongest fixed baseline is 0.141 percentage points; the gap to the oracle is 0.140 points.

The oracle is a hindsight upper bound and has access to measured unguarded or frontier-guarded variants as well. The deployable selector chooses a rate-guarded representative. This distinction matters when interpreting the remaining gap.

Where the rule fails is part of the result.

All four family-selection errors occur on the Tencent traces: 8/9 and 6/9 settings correct. Each of the other traces has 9/9. Those related settings should not be presented as independent deployments; there are seven trace-level holdouts.

A follow-up using collected browser-agent traffic changes both the workload construction and the winning policy families. Some settings favor ARC, which is outside the original selector’s action space. That is a reason to revisit the method’s coverage.

Synthetic-to-real transfer remains unproven. The follow-up also changes feature extraction and policy coverage. Its results are not a controlled before/after score for this exact experiment. The original 93.7% applies to this synthetic-overlay panel.

The design lesson is to track the objective directly. Family accuracy diagnoses the selector, while hit rate measures the effect we wanted from it. A useful research artifact should make both visible, including the settings that break the rule.

Reproduce the evaluation

The standalone release contains aggregate policy measurements, frozen prefix features and the selector implementation. No private research-workspace history or raw request logs are included.

git clone https://github.com/landigf/cache-regime-benchmark.git
cd cache-regime-benchmark
python3 scripts/reproduce.py

This checks input hashes, reconstructs winners from the archived policy cells, refits every held-out fold and compares the predictions against the archived record. It requires Python 3.10+ and the standard library. It does not rerun raw-trace collection or simulation.

Methods & limitations · Recomputed JSON · Data sources & terms