7 tracesone trace held out for each evaluation fold
+0.141 ppmean hit-rate gain over the best fixed baseline
A selector is useful only if the cache benefits.
Different workloads reward different eviction policies. I studied whether a short warm-up prefix contains enough information to choose a useful policy family before replaying the rest of a mixed workload.
The experiment combines seven public cache traces with a WebLINX-derived synthetic agent overlay. Each trace is evaluated under three overlay behaviors and three target agent fractions, giving 63 settings at a 64 MiB cache size. The selector profiles the first 1,000 mixed requests. It chooses a rate-guarded representative from WTinyLFU, S3FIFO or GDSF.
The evaluation holds out an entire trace, including all its settings. A small threshold tree is fitted on the other six traces. This asks whether the rule transfers to a different trace within the panel.
Recomputed from 513 archived policy measurements and frozen prefix features. Rates are unweighted means over the 63 settings. Agent traffic is a synthetic overlay; the cache budget is 64 MiB.
Mean human-request hit rate. “Human” denotes the original public-trace component in the mixed workload.
Approach
Hit rate
How the policy is chosen
Per-setting oracle
13.22%
Best measured policy after seeing the outcome
Held-out selector
13.09%
Tree trained on the other six traces
Fixed WTinyLFU + rate guard
12.94%
Same policy in every setting
Fixed S3FIFO + rate guard
11.77%
Same policy in every setting
Fixed GDSF + rate guard
10.67%
Same policy in every setting
The selector matches the winning family in 59 of 63 settings, or 93.7%. That accuracy does not translate into a 93.7% cache hit rate. The practical gain over the strongest fixed baseline is 0.141 percentage points; the gap to the oracle is 0.140 points.
The oracle is a hindsight upper bound and has access to measured unguarded or frontier-guarded variants as well. The deployable selector chooses a rate-guarded representative. This distinction matters when interpreting the remaining gap.
Where the rule fails is part of the result.
All four family-selection errors occur on the Tencent traces: 8/9 and 6/9 settings correct. Each of the other traces has 9/9. Those related settings should not be presented as independent deployments; there are seven trace-level holdouts.
A follow-up using collected browser-agent traffic changes both the workload construction and the winning policy families. Some settings favor ARC, which is outside the original selector’s action space. That is a reason to revisit the method’s coverage.
Synthetic-to-real transfer remains unproven. The follow-up also changes feature extraction and policy coverage. Its results are not a controlled before/after score for this exact experiment. The original 93.7% applies to this synthetic-overlay panel.
The design lesson is to track the objective directly. Family accuracy diagnoses the selector, while hit rate measures the effect we wanted from it. A useful research artifact should make both visible, including the settings that break the rule.
Reproduce the evaluation
The standalone release contains aggregate policy measurements, frozen prefix features and the selector implementation. No private research-workspace history or raw request logs are included.
git clone https://github.com/landigf/cache-regime-benchmark.git
cd cache-regime-benchmark
python3 scripts/reproduce.py
This checks input hashes, reconstructs winners from the archived policy cells, refits every held-out fold and compares the predictions against the archived record. It requires Python 3.10+ and the standard library. It does not rerun raw-trace collection or simulation.