Notes from working with generative AIs — told through the science fiction I read as a boy.
Four stories, four lessons: Sturgeon’s creator of a fast little civilization,
and Niven’s pilot who is the slowest part of a very fast ship.
Introduction2026-09-03
What this is — and what’s coming
I’m writing up some of my experience using genAIs.
And the more I use them, the more I am reminded of the science fiction I read as a boy.
Two parts.
First, the results — the deployed ENS_OPT momentum strategy: the ladder of what each risk layer adds, what didn’t work, growth of $1, and the exposure dashboard.
Then, the lessons — told through the science fiction: what the agents are, and what I become when I work with them.
Results2026-09-03
The ladder layers, explained
N100 [3,11] top-4 leg
The engine. Each month, rank the Nasdaq-100 by momentum measured
over a blended 3- and 11-month lookback and hold the top four names, rebalanced monthly. This is the
raw momentum effect every column is built on.
AbsMom5 timer (N100 leg)
A binary trend filter: when the index's own trend turns down
the N100 leg steps entirely to T-bills, otherwise it is fully invested. It rescues sustained bear
markets (see the 2000s) but whipsaws in choppy drop-then-rally regimes (see the 2020s), because it
exits after a fall and re-enters after the recovery.
SP100 [14] top-1 leg @ 40%
A second, diversifying sleeve — the single strongest
S&P-100 name on a 14-month lookback, held untimed, at 40% of the portfolio. The two sleeves
rarely bottom together, so the blend roughly halves the worst drawdown versus either alone, at the
cost of some upside in decades the Nasdaq dominates.
Geo de-concentration [1,2,4,8]
Instead of weighting the four names equally, weight them
1:2:4:8 by rank so the lower-ranked, less-crowded names carry more of the leg. Spreading away from
the single hottest name modestly improves risk-adjusted return.
Vol dial (target 25%)
A continuous volatility-targeting overlay: raise exposure when
recent volatility is low, cut it when high, aiming for ~25% annualized volatility. Formally the
month's exposure is w = min(25% ÷ σ, cap), where σ is the strategy's trailing
12-month annualized volatility (lagged one month, so the weight is causal) and cap =
1.5× for Growth or 1.0× for Growth (no margin); the invested fraction earns the strategy, the
uninvested fraction earns T-bills, and any borrowed fraction (w > 1) pays cash plus
a spread. Both Growth columns use it; they differ only in whether leverage is allowed (see
Margin).
Margin (gross exposure > 100%, up to 1.5×)
The full Growth dial lifts gross
exposure above 100% in calm regimes — up to 1.5× — which requires margin
(borrowing, financed at the cash rate plus a spread in this model). Growth (no margin)
forbids this: it caps exposure at 100%, so it never invests more than the current balance —
keeping the dial's downside protection without leverage, and giving up the calm-regime return boost.
On the stacked panel below, bands past the dashed 100% line are levered months; they occur only for
the margined Growth column.
Results2026-09-03
The strategy ladder — what each layer adds
Layer / metric
Base
Original
Blend eq N100 + SP100 · no geo
Aggressive deployed · + geo
Growth (no margin) experimental · unlevered
Growth experimental · margin
N100 [3,11] top-4 leg
✓
✓
✓
✓
✓
✓
AbsMom5 timer (N100 leg)
—
✓
✓
✓
✓
✓
SP100 [14] top-1 leg @ 40%
—
—
✓
✓
✓
✓
Geo de-concentration [1,2,4,8]
—
—
—
✓
✓
✓
Vol dial (target 25%)
—
—
—
—
✓
✓
Margin (gross > 100%, up to 1.5×)
—
—
—
—
—
✓
CAGR (annualized)
2000s CAGR
6.2%
25.4%
25.9%
29.6%
28.3%
31.6%
2010s CAGR
38.0%
38.2%
26.8%
27.3%
27.2%
38.3%
2020s CAGR
55.0%
34.1%
54.8%
54.1%
46.7%
47.8%
Total CAGR since 2000
28.8%
32.3%
32.9%
34.4%
32.2%
38.0%
Returns
2000s return
1.8×
9.5×
9.8×
13.1×
11.8×
15.3×
2010s return
25.1×
25.3×
10.7×
11.2×
11.1×
25.6×
2020s return
17.9×
6.9×
17.8×
17.3×
12.4×
13.1×
Total return since 2000
808×
1,654×
1,877×
2,531×
1,636×
5,119×
Drawdowns
2000s drawdown
-72.3%
-32.8%
-34.1%
-32.2%
-30.0%
-30.0%
2010s drawdown
-17.1%
-18.8%
-22.3%
-26.3%
-24.1%
-24.7%
2020s drawdown
-49.0%
-49.8%
-30.3%
-29.3%
-15.8%
-19.7%
Total drawdown since 2000
-72.3%
-49.8%
-34.1%
-32.2%
-30.0%
-30.0%
Sharpe (risk-adjusted)
2000s Sharpe
0.35
1.00
0.94
0.98
1.03
1.03
2010s Sharpe
1.48
1.64
1.27
1.27
1.29
1.32
2020s Sharpe
1.16
0.91
1.27
1.29
1.38
1.34
Total Sharpe since 2000
0.90
1.11
1.10
1.13
1.19
1.21
Common window 2000 → 2026-07. Aggressive is the deployed strategy. All figures hypothetical, frictionless — not advice.
Results2026-09-03
What didn’t work (1 of 2)
The other half of an honest record — improvements tried and set aside.
Timing the SP100 leg. Blending with both legs timed (a retired 45/55 weighting)
gave a lower Sharpe than the N100 sleeve alone under the live configuration — timing the second leg
subtracted value. The SP100 leg is kept untimed and the blend reweighted to 60/40.
Rebalancing more often than monthly. Weekly and biweekly clocks lost on every metric
before costs and deepened drawdowns; the extra trades chase days-scale noise. Momentum's timescale
is months.
Reweighting or changing the momentum lookback. No lookback variant robustly beat the
equal-weight [3,11]; the two that beat it over the full period both failed a fixed-baseline test
across market sub-regimes that [3,11] passed.
Alternative ranking signals. Regression slope (274- and 366-day), fitted alpha
(252-day), the K-ratio, and OLS-alpha were each tested as the ranking signal; none beat a simple
trailing-return ranking over the long run.
Chasing last month's winner. Ranking on the most recent month alone underperforms — the
one-month winner is anti-persistent, tending to mean-revert over the following month. This is why
the signal blends multi-month lookbacks (the [3,11] window) rather than the latest month.
Results2026-09-03
What didn’t work (2 of 2)
Extending to sector and single-country ETF universes. Momentum is real there but at
roughly a third of the core magnitude, absent through the 2010s, and a buy-and-hold control showed
the blend's apparent benefit was mostly market beta. Left out of the run list.
Gap-up / event momentum in large caps. A survivorship-bias-free replication of the
"Power of Price Action Reading" gap strategy across 28 universes found the effect strong in
micro-cap, small-cap and biotech but weakest in the mega-cap Nasdaq-100 / S&P-100 names this
strategy trades (~+0.3R vs +0.6R at 30 days), and largely a 2020 artifact there. Liquid mega-caps
absorb gaps too fast, so folding gap/catalyst entries into the deployed universe imports the
thinnest version of the effect. Kept as a separate study.
Concentrating the portfolio. Overweighting the single top-ranked name was tested and
rejected; the deployed allocation instead spreads weight toward the lower-ranked names via geometric
de-concentration [1,2,4,8] (linear [1,2,3,4] is a milder variant). This is an ensemble-level, timed-
blend effect — on the raw untimed sleeve alone, equal weight actually edges it out — so it is applied
within the deployed strategy, not as a universal rule.
The vol dial as a deployed feature. It lowers drawdown but gives up raw return unless it
can use leverage, and that leverage benefit depends on calm regimes that may not recur. Kept
experimental, not deployed.
Results2026-09-03
Growth of $1 (Jan 2000 →), log scale
The two dashed lines are a
separate strategy, not a ladder layer: Catalyst
— the uncorrelated macro-regime satellite (data starts 2003) — and a
33% Catalyst / 67% ENS_OPT blend, both anchored to the
deployed curve at Catalyst’s 2003 start so they read on the same scale. Catalyst trails on raw growth
by design; the blend’s payoff is risk-adjusted (roughly half the drawdown at a similar path) because
the two barely correlate — see the Catalyst study
(ρ 0.21). This is a blend, not a stacked ladder layer, which is why it lives on the chart
and not in the ladder table above.
Results2026-09-03
The exposure dashboard
Three shared-time panels: trailing volatility versus the 25% target (with holdings correlation); the stacked holdings by weight against the 100% line; and the equity curves — deployed (Aggressive) versus the Growth dial.
Results2026-09-03
ENS_OPT monthly timeline — deployed 60/40
The exposure graphic, month by month: the four N100 holdings (by rank weight) and the SP100 name, the AbsMom5 timer, the per-month exposure bar, and the return — deployed versus the Growth dial.
Lessons2026-09-03
Lessons
What the work taught me — told through the science fiction I read as a boy.
Two old stories describe the two halves of the experience:
what the agents are, and what I become when I work with them.
Lessons · The story behind the title2026-09-03
Theodore Sturgeon’s “Microcosmic God”
James Kidder is a brilliant scientist living on an isolated island.
He realizes he lacks true creative genius for broad innovation, but excels at perfecting
and building practical applications from ideas.
Impatient with slow human progress, Kidder creates a synthetic, rapidly evolving race of tiny
humanoid creatures — the Neoterics.
With an accelerated metabolism and lifespan, the Neoterics develop a high-tech civilization in
days or weeks — advanced gadgets and inventions, under Kidder’s strict rules and demands.
The parallel
My genAI agents are Neoterics: a fast little civilization that invents under my rules.
The rest of this talk is what I’ve learned about keeping them honest — and about who the slow one is.
Lessons · The experiment, honestly2026-09-03
What actually happened — a corrected summary
I had the AI read my momentum code and write an article about it.
Then I asked it for ways to improve CAGR and drawdown.
That became a long test-and-cull loop: keep what survives an out-of-sample check,
discard the rest — it said “no” more often than “yes.”
A real chunk of time went to CIMI ideas (slope, walk-forward) — the same skepticism,
pointed at someone else’s results.
The correction
The keep/discard calls weren’t the model’s taste — they were driven by backtests,
out-of-sample skepticism, and my judgment. The value wasn’t a genius picking winners; it was cheap,
fast hypotheses paired with a disciplined filter.
Claude’s point of view: Bill used the LLM as a tireless
idea-and-implementation engine, then refused to trust any of it until it survived an out-of-sample,
adversarial check — several of which the LLM itself failed. The moat wasn’t the model’s
cleverness; it was the redundancy around it.
Lessons · The loop had teeth2026-09-03
What the filter kept — and what it culled
Kept — survived the check
Cash earns the 13-week T-bill, not 0%.
Allocation de-concentration — overweight the lower-ranked names.
A volatility-targeting overlay — scale exposure to recent volatility.
Culled — failed the check
Weekly / biweekly rebalancing — monthly wins before costs.
Lookback-weight reshuffling — no robust gain over [3,11].
Sector-rotation filter — fine alone, negative when blended in.
Country / sector-ETF universes — real momentum, wrong magnitude.
Strat-of-strats / walk-forward selection.
Claude’s CIMI opinion Slope: the implementation was correct, but it never beat plain [3,11].
Walk-forward: impressive-looking, but it raises survivorship and
multiple-testing questions I’d want answered before adopting it.
Lesson 1 · redundancy2026-09-03
My Finance Neoteric’s cross-check
Errors occur in my trading-system work. Occasionally the data is bad (e.g. NaN prices);
occasionally the agent makes mistakes.
How did I discover this? As I created more artifacts, some headline numbers — e.g. total
return over 20 years for the same strategy — differed between artifacts.
I noted it (or the AI caught it itself) and reported it.
One way I’m avoiding it: have the agent build big spreadsheets that reproduce the calculations.
If the AI’s answers don’t match the spreadsheet, I ask “why?”
Lesson
Redundancy. One big checksum on all the work.
Lesson 1 · redundancy (the war story)2026-09-03
…and here is the time it caught me
The first article I had the AI write about my strategy was flattering — and partly false.
The published claim — later retracted “The strategy completely avoided the 2022 growth-stock carnage… the timing signal spent 2022 preserving capital.” Reality: it returned −36.8% in 2022, inside its worst-ever drawdown —
−49.8% (Oct 2021 → Apr 2023), not recovered until Oct 2025.
A later audit found ten errors across the published articles and deck — and
none of them were arithmetic. Every one was a true statement about one configuration,
applied to another: a timed leg described as untimed, an equal-weight figure quoted as geometric,
a 45/55 blend reported as 60/40.
Confident, fluent, wrong — the same failure mode as the flattering first draft.
So I built a claims register: it re-checks the foundational facts on every change and
exits with an error if anything is refuted. Today: 49 facts verified, 0 refuted.
Lesson
The checksum isn’t paranoia — it already caught me. Fluent and confident is not the same as correct.
Lesson 2 · separate processes2026-09-03
My Medical Neoteric’s workflow
Product Marketing
spec
——→
test result
←——
Product Development
Sometimes it makes sense to use multiple agents — and I try to leverage the
inter-agent competition I’ve noticed.
For my patient-tools work: a product-marketing agent writes a spec; a development agent
implements it and gives feedback; then marketing tests it and fills gaps in a new version of the spec.
Specs and feedback are versioned — a history of decisions.
No QA agent: marketeer and programmer build the test plan between themselves (it’s part of the spec).
Marketing wants to control development directly — probably a mistake — so I’m weighing an
orchestration agent (COO) to watch both. I fill that role now; it’s time-consuming.
Marketing also makes unjustified assumptions about the world, so I may add a strategic marketer
to go gather real information. I fill that role too.
Lessons · What changed my mind2026-09-03
Super Agent vs. Super Agents
My finance and medical exercises changed my opinion about one thing.
Before: as a one-man consultancy who did everything, I felt strategic marketing, product
marketing, development, operations, etc. were artificial divisions — and that genAI let me
build a new, less “human-centric” (and likely better) organization.
Now: the artifacts’ chief function (specs, product-as-implemented, test sets) is to
enable error correction — and the functional roles exist to produce those artifacts.
Why not produce spec and implementation with one agent?
When I collapse the roles, the false assumptions the agent held while writing the spec carry
straight over into the implementation.
So specs alone are not enough. I need separate processes with separate memories.
I suspect agents differ because each randomly initializes to a different state.
Lessons · The second story2026-09-03
Larry Niven’s “At the Core”
The ship: Beowulf Shaeffer is hired by a Pierson’s Puppeteer named Nessus to pilot the
Long Shot, a prototype equipped with a radically fast new quantum hyperdrive.
The problem: the ship travels so fast — light-years in minutes — that the stars ahead
blur into a solid wall of light. Shaeffer must repeatedly drop out of hyperspace just to check his
position and make sure he isn’t about to crash into a star.
The conflict: when Shaeffer suggests flying above the plane of the galactic disk, where space is
empty, Nessus refuses. The Puppeteers insist he stick to the pre-planned, commercial flight path through
the dense star clusters — because the entire trip is a marketing stunt to sell the hyperdrive.
Lesson 3 · the human2026-09-03
I am the bottleneck
Like Beowulf Shaeffer, I move slow, but the ship moves fast.
My day is spent moving from one window to another — saying “Yes,” or reading answers to
my questions.
This happens even when I use multiple agents to work independently on the same problem.
Lesson
The bottleneck isn’t the drive. It’s the pilot who has to keep dropping out of hyperspace to check the position.
Lesson 4 · structure2026-09-03
The power of telling a good joke
One other thing I’ve noticed: when communicating with LLMs, it helps to be able to tell a good,
structured story.
I’m positive this has a lot to do with the next-word-prediction approach LLMs take to language.
Think about punch-line jokes especially: if the setup is clear and well structured, there are only a
finite number of ways to complete the narrative — unexpected, perhaps, but finite.
Lesson
Programming LLMs well is a lot like telling a joke where the LLM gets to step on the punchline.
Lessons · An LLM-style joke · three completions2026-09-03
…three ways to finish the story
A bird was flying south for the winter. Late in the season, it froze up and fell into a field.
To add insult to injury, a cow came over and shat on him. But the dung was warm, and the bird began to thaw.
When it realized it wasn’t going to die, it poked its head out of the pile and began to sing.
A cat heard the singing, came over, pulled the bird out… and ate him.
There are three morals to this story
Not everyone who shits on you is your enemy.
Not everyone who gets you out of the shit is your friend.
And if you’re sitting in shit and happy — don’t sing.
The setup constrains the endings. That is exactly the leverage you have with an LLM.
Other items2026-09-03
Other items
Fundamental checks
Test set
Prompt library
Data source — Renaissance Capital: reduce variance either by trading a lot, or by getting more data.