# Before and after training: food-response extension

Kuber Mehta · 11 September 2026. Proposed extension, not a frozen experimental
protocol or a completed result. The registered BabyLM comparison keeps its
training priority after the completed language computation controls. Existing studies,
conditions and tests are unchanged.

The [six-channel sensory interface](food-sensor-interface.md) is now implemented
and calibrated against actual body positions in two short prescribed-motion
FlyGym repetitions. It separates bilateral odor, sugar contact and source
contact, with an independent scalar geometry audit. A separate
[scripted odor-approach reference](food-approach-results.md) now reaches the
odor-A patch in the two declared mirrored situations, ignores sugar reversal,
and times out when odor is removed. Its seven physical trials use no neural
policy or training. Learned food behavior, language transfer and feeding remain
untested. Four paired initial/language-trained WikiText cores now pass
[numerical sensory replay](food-core-interface.md) through identical fresh
adapters, with no food adaptation or physical feedback. The complete-group
selection comparison remains next after the fixed BabyLM queue.

The subsequent [24-case physical reference](food-core-physical-results.md) is
complete and independently replay-audited across 4,824 observations. Both initial
and language-trained cores contact the upper patch in mirrored cases, including
wrong-source contacts. All conditions now have synchronized recorded neural
state and reconstructed body geometry in the browser, with the
[body-state limitations](physical-state-semantics.md) explicit. This is before
food adaptation; the downstream learning and transfer experiment below remains
unrun.

An [episodic reward-modulated action readout](food-readout-learning.md) is now
implemented with mathematical and resume checks. Its first scope freezes both
the sensory projection and recurrent core, learning only the 771-entry action
head. This narrows the trainable interface explicitly; it supplies no physical
adaptation result and does not freeze the pending experimental budget.

The [episode runner](food-episode-runner.md) now connects sampled readout actions
to the existing physical environment. Eight short, repeated integration checks
ran with learning disabled; body and neural arrays repeated exactly. This
supplies no navigation, adaptation or language-benefit result. Full-horizon
sampled-policy references remain pending. A [paired schedule executor](food-adaptation-schedule.md)
now checks complete training inventories, readout chaining and evaluation
isolation with toy transitions. Its official registration, cost/prerequisite
checks, physical schedule and analysis remain undeclared.

## What can be compared now

The browser's physical-choice replay compares the original untrained sensory
network with checkpoints at updates 300, 600 and 900. All five learning methods,
both cues and all 40 recorded trials remain available. The trained rule reverses
between updates 301–600 and returns between 601–900; the initial reference is
evaluated under the original rule. A changed direction can therefore be correct
or incorrect depending on the task phase.

The two panels show recorded thorax position and heading at matching times and
with identical coordinates. They use every tenth stored frame without
interpolation. The first stored observation is at 0.0001 s, not an invented
reset frame at zero. The fly icons are schematic and are not reconstructed
joint motion. All 40 trials produce only two distinct physical paths: the
network selects one of two designed motor commands. This is one training seed
and one physics seed, not 40 independent behavioral replications.

The [exporter](choice_replay.py) checks all original trajectory hashes,
complete condition coverage, greedy neural choices, shared initial responses,
and exact equality of recorded time, position, heading and generalized position
arrays within each command before sharing those two paths. The
[replay data](choice-replay.json) retain every checkpoint and
trajectory identity. Original states and physical outcomes remain in the
[complete records](learned-choice-records.json).

These are abstract sensory cues. They have not been replaced by fruit, sugar,
language-trained weights or a learned gait. The typing fly in ChatFLM is a
separate illustration driven only by generation status.

## Separate the two food questions

**Finding a fruit-like source:** use a declared odor field sampled at the body's
sensor locations. Compare arrival, search paths and retention after association
learning. NeuroMechFly v2 already supports olfaction and learned navigation;
that is platform precedent, not a new FLM capability or result.
[Wang-Chen et al., 2024](https://www.nature.com/articles/s41592-024-02497-y).

**Responding to sugar contact:** introduce a separate contact/taste input and
measure a declared action or neural readout after contact. Do not treat a sugar
label as an odor sensor or equate arrival at a waypoint with feeding. Prior
whole-brain work stimulated identified gustatory neurons and tested predicted
feeding circuitry, including motor-neuron responses. That circuit-specific
evidence cannot be transferred to this small, differently selected language
subset. [Shiu et al., 2024](https://www.nature.com/articles/s41586-024-07763-9).

The first FLM extension should be described as an engineered food-association
task. It needs explicit sensor units, action meaning, reward timing, reset rules
and contact geometry. A biological feeding claim additionally needs verified
sensory/motor neuron identities, circuit coverage and an independently validated
feeding readout. The current [subset audit](subset-audit.md) does not establish
that coverage; rendered proboscis motion would not supply it.

## Comparison to declare before running

Use the same sensory and action interface, initialization seeds, environments,
downstream exposure and evaluation layouts for the original random core and
its paired language-trained core. Evaluate before any task adaptation and at
fixed downstream checkpoints. Freeze the core in the primary transfer comparison
so learning the new interface is distinguished from changing recurrence again.
Declare any trainable-core extension separately. A trained sensory-only model
provides a task-learning control, not a language-transfer substitute.

Include a scripted sensor-to-action reference, fixed/random recurrence and the
matched rewired language cores when making an anatomical claim. Match sensory
exposure and adaptation settings within each paired comparison. Keep lateral
recurrence, temporal state and learned readout contributions distinct; preserve
their parameter and compute accounting. Pretrained groups have additional
lifetime language exposure, even when downstream exposure matches.

Counterbalance which source predicts reward. Reserve unseen layouts and odor
fields for evaluation, and include neutral sources, missing-cue trials, reward
omission and a declared reversal/retention phase. Reward history and task phase
must not leak through an unreported input. Record every success, wrong approach,
timeout and simulation failure. Predefine arrival/contact thresholds and a
primary measure, such as correct-source arrivals within the time budget; retain
latencies with failures marked as censored rather than averaging only successes.

Compare the paired change from before to after task adaptation, then compare
that change across initial and language-trained cores. Report the unchanged,
beneficial and harmful cases. At best this would establish transfer through the
declared interface; it would not show that text gave an animal biological
knowledge of food. Seeds, training budgets, stimulus files, checkpoint identities
and analysis must be frozen before fitting or selecting demonstrations.

## Delivery after that experiment

Add synchronized before/after physical recordings beside measured action
probabilities, task outcomes and actual recurrent activity. Identify whether
the displayed difference follows language pretraining, food-task adaptation,
both or neither. Keep the complete condition selector and downloadable records,
including cases where the fly does not improve. The page now displays the
existing cue-learning recordings and the unadapted physical reference; no
post-food-adaptation result is available yet.
