A good PQ cue should add pronunciation information while producing the smallest possible disturbance to the visual information the reader needs to identify the letter, its position, and the recurring multi-letter pattern.
That gives us a way to evaluate the existing PQ cue vocabulary.
First, what the evidence says we should protect
Learning to read actually changes ventral visual cortex: word selectivity emerges with reading acquisition, and print sensitivity appears very early as children learn letter–speech-sound correspondences. Skilled reading is associated with increasingly tuned word-selective responses.
But the relevant visual representation isn’t merely “I recognize these individual letters.” Multi-letter visual structure matters. In children, experimentally disrupting multi-letter features with mixed case slowed learning and subsequently impaired recognition of the trained forms, even though single-letter naming was unaffected.
And letter recognition itself depends upon spatial features. Features can be masked or even perceptually migrate between neighboring letters, while crowding is a major limitation on how many letters can be accurately apprehended around fixation.
That suggests three things PQs should preserve simultaneously:
letter identity → letter position → multi-letter configuration.
The PQ cue should preferably modify a secondary visual dimension while leaving those three as intact as possible.
Applying that criterion to PQ cue types
| PQ visual operation | Likely compatibility with visual-word learning | What I would test |
|---|---|---|
| Gray / reduced contrast for silent letters |
Very promising | How faint can it become before letter identity/word configuration suffers? |
| Moderate bolding for letter-name sound |
Promising | Minimum weight change that is reliably discriminable |
| Underline / connection for cooperating letters |
Promising, if subtle | Whether it increases grouping without masking descenders/features |
| Dotted underline / alternate connection | Promising | Whether children reliably distinguish it from ordinary underline |
| Small vertical displacement | Use cautiously | Recognition and word-shape disruption thresholds |
| Horizontal stretching | Use cautiously | Whether altered width disrupts neighboring-letter localization |
| Rotation | Most important to test carefully | Maximum rotation that preserves immediate letter identity |
| Kerning/spacing manipulation | Highest concern | Whether it disrupts letter position, grouping or visual span |
| Inserted segmentation marks | Potentially useful but intrusive | Whether segmentation benefit exceeds disruption of recurring word form |
Those are hypotheses from the vision literature, not findings about PQs themselves.
Gray/silent letters
This may be one of the more elegant PQ operations.
The letter remains:
- the same letter,
- in the same position,
- at essentially the same size,
- with the word’s horizontal geometry intact.
What changes is primarily contrast/salience.
That’s almost an ideal information channel for something like: “recognize that this letter exists orthographically, but don’t give it an independent sound.”
Importantly, I would not make silent letters extremely faint. PQs want the learner simultaneously learning the actual spelling. If
e
effectively disappears perceptually, we’ve started teaching an altered word.
The optimization target should therefore be:
maximally perceptible as an e while minimally competitive as a sound-bearing element.
That’s experimentally measurable.
Bold for letter-name sound
There is surprisingly relevant evidence here.
One lexical-decision study found bold words were recognized faster than ordinary words, specifically for lower-frequency words. But bold isn’t intrinsically beneficial: extreme stroke weights can reduce letter recognition, and a 2025 eye-movement study found that mechanically bolding portions of every word—the “Bionic Reading” approach—could actually impose reading costs.
That’s actually encouraging for the PQ use of bold, because PQ bolding isn’t decorative or mechanically positional.
Its occurrence carries information:
this letter is producing its letter-name sound.
Consequently, the child can potentially learn a three-way co-implication:
visual identity of I ↔ salient/bold I ↔ auditory /aɪ/
while ordinary
i
remains visually
i
.
I would therefore preserve bold but use a
moderate rather than extreme weight differential.
Width/stretching deserves more thought than I initially gave it
There is evidence that wider letter shapes can actually improve parafoveal/peripheral recognition compared with narrow
ones. So horizontal expansion isn’t inherently hostile to letter recognition.
But PQs create a different situation: one letter within a word changes width.
That changes the spatial relationship among letters. Because feature location and crowding matter, this could create
either a benefit or a cost.
Therefore I wouldn’t eliminate stretching. I’d ask:
Can the stretch be large enough to communicate the PQ distinction while small enough that the visual system
continues to treat it as an ordinary instance of that letter occupying its ordinary serial position?
That’s precisely the sort of parameter PQ testing should determine.
Rotation is particularly interesting
Here we need to separate two things. The visual system is remarkably capable of recognizing objects across transformations. But letters are a peculiar
category because orientation can itself carry identity
—
b/d
,
p/q
, etc.
So your rotated-R convention has a potentially useful property and a potential cost. It can make an unusual sound relationship perceptually conspicuous while preserving substantial R structure. But it alters a geometrical feature that the developing visual system may use for letter recognition.
I would therefore not redesign the existing rotated R from theory alone . Instead, parameterize its existing angle and experimentally find the lowest rotation at which children reliably
detect the PQ distinction.
The goal isn’t:
make the cue obvious.
It’s:
make the cue just obvious enough.
That’s a recurring theme across these dimensions.
Underlining/grouping may be especially compatible with what PQs is trying to do
If two letters cooperate in producing a sound, connecting them without altering their actual internal forms has an attractive property:
the informational modification exists between the letters rather than inside them.
th
can remain visually normal
t
+
h
, while an external relational mark says:
treat these together. That potentially preserves letter-feature recognition while adding information about higher-order grouping.
And this is precisely where the finding about multiletter features becomes relevant: developing readers aren’t merely accumulating independent letter identities; recurring multi-letter visual structures contribute to orthographic learning. So relational cues may be particularly valuable when the PQ information itself concerns a relationship among letters. There’s almost a design principle hiding here:
Encode information about a letter by minimally changing the letter. Encode information about relationships among
letters by modifying the perceptual relationship among them.
That seems deeply compatible with PQs.
Spacing/kerning is where I would be most conservative
Crowding and relative position are fundamental to reading. Feature migration between adjacent letters occurs, and crowding substantially limits the visual span available during reading.
Changing spacing therefore modifies a dimension already doing important work.
This doesn’t mean PQ spacing cues are bad. Increased separation could sometimes reduce crowding. But it could simultaneously weaken perception of the recurring letter cluster as a unit.
So I would avoid large spacing manipulations unless they carry information that can’t be communicated through a less structurally consequential channel.
But I think there’s a deeper implication for PQ design
We have been thinking about PQs as if each cue represents a sound distinction.
VWFA/perceptual-learning research suggests another way to formulate them.
Imagine three layers:
Invariant
e
=
this is the letter e
Contextually variant
e
=
this particular e is participating in this particular sound relationship
Recurring experience
many similarly cued
e
s across many words =
the learner discovers what this visual/sound contingency predicts
That last level is crucial.
The learner doesn’t necessarily have to consciously formulate:
“When E occurs in configuration X under phonological condition Y, apply rule Z.”
The perceptual system is extraordinarily good at learning recurring feature contingencies. Experiments on letter strings specifically find that letter processing is contingent on neighboring features rather than completely independent.
So PQs could be understood as increasing the signal-to-ambiguity ratio of the statistical environment from which orthographic learning occurs.
Ordinary English gives the learner something approximately like:
same visual feature → A, B, C, D… possible auditory consequences.
PQs temporarily provide:
visually differentiated variant A → auditory A
visually differentiated variant B → auditory B
visually differentiated variant C → auditory C
But critically, all three remain recognizably instances of the same underlying conventional letter.
As learning progresses, the surrounding orthographic contexts themselves become increasingly predictive. The added PQ feature therefore becomes informationally redundant.
That gives us a much more precise conception of cue fading:
Don’t fade PQs merely because the child has reached a grade or completed N exposures. Fade a cue when the learner
demonstrates that the uncued orthographic context has acquired the predictive function previously supplied by the PQ feature.
That would make adaptive PQ fading itself a measure of orthographic learning.
And that suggests a concrete PQ research program
Rather than asking merely, “Do children read faster with PQs?”, I would construct a feature-transfer experiment.
Teach novel or initially difficult words under several conditions: ordinary orthography, full PQ, and selectively reduced PQ. Then test them later only in ordinary orthography. Measure accuracy, latency/fluency, eye movements if feasible, and transfer to untrained words sharing the same spelling–sound pattern.
The decisive PQ prediction isn’t simply: PQ words are easier while PQs are visible.
It is: PQ-supported encounters produce faster acquisition of accurate recognition of the corresponding uncued orthographic
patterns.
If that happens—and particularly if children generalize the relationship to untrained words containing recurring spelling patterns —we would have evidence for something much closer to your original conception: PQs aren’t functioning primarily as an assistive overlay. They are
temporarily differentiating the perceptual learning environment so that the learner can more readily discover the recurring relations already latent within conventional orthography.
And that, rather than VWFA activation itself, is the claim I would put at the center of a serious PQ/VWFA research proposal.
Review: Reading — The Confluence of Vision and Language
Study: Brain sensitivity to print emerges with letter–speech-sound learning
Are you being careful to avoid circular logic here? VWFA is not an innate naturally occurring brain organ. Its an abstract observational conceptual overlay. Whatever the actual neurology, it is a result of learning orthography in the ways it could only be learned given the static invariant characters used to learn it with. In other words, it reflects the way we have learned not necessarily the optimal way to learn. On the other hand we know, and it’s buttressed by eye tracking and other studies, that the more confusing the letter sound correspondences, the greater the processing overhead that leads to stutters that erode attention span and disentrain comprehension.
Yes. Your correction exposes a circularity risk in the way I framed the previous answer.
I was partly reasoning from what the conventionally trained reading system becomes back toward how PQs ought to train it. That inference is not warranted. The VWFA literature is much more useful descriptively than normatively.
The evidence itself supports your distinction. Word-selective responses at the so-called VWFA site are absent before reading and emerge with literacy acquisition; longitudinal studies show the change occurring as children acquire reading, and experimental reading instruction changes responses in visual cortex. Indeed, there has long been disagreement about whether “VWFA” should even be understood as a discrete specialized word-form module rather than a learned pattern of participation by ventral visual cortex in a distributed system.
So I would restate the relationship much more carefully.
VWFA tells us what learning conventional orthography does, not what reading has to be
The observed system has developed under a peculiar training regime: static invariant conventional characters + highly variable spelling→sound correspondences + enormous repetition →eventual fluent recognition.
The resulting neural organization therefore contains the history of adaptation to that environment.
It cannot, by itself, tell us that preserving exactly those visual conditions is optimal during acquisition.
In fact, neuronal-recycling accounts explicitly treat literacy as cultural learning recruiting pre-existing visual circuitry rather than uncovering a biologically pre-specified reading organ. A very recent longitudinal study likewise describes literacy acquisition and experience-dependent connectivity as driving VWFA specialization.
That changes the PQ question substantially.
The stronger evidence for PQ’s premise comes from the cost of ambiguity
This is where orthographic-depth, behavioral and eye-movement evidence becomes more important than VWFA localization.
English is a deep orthography: grapheme–phoneme relations are many-to-many, context dependent, and affected by multi-letter and morphological structures. Cross-linguistic evidence shows that orthographic depth materially changes reading acquisition and the strategies learners employ.
Eye tracking also shows consequences at the level you’re concerned with. German and English readers presented with closely matched materials exhibit different word-recognition dynamics attributable to orthographic consistency; effects remain detectable even in skilled adults. Orthographic familiarity itself affects eye-movement behavior, demonstrating that properties of the orthographic signal influence ongoing allocation of visual processing during reading.
And this is important for your comprehensional stutter idea: during normal reading, phonological information is activated extremely rapidly—within the time course of a
fixation—rather than being some leisurely subsequent process. So the more defensible causal chain is:
orthographic ambiguity
→ greater uncertainty/competition in resolving print to familiar language
→ additional processing time/effort and/or rereading
→ interruption of fluent word-to-word processing
The next step you propose—
repeated interruptions erode attentional continuity and therefore disentrain the ongoing comprehensional process
—is plausible and, I think, extremely important, but I would not yet call the whole chain experimentally established. The evidence is considerably stronger for the earlier links than for the final “disentrainment” construct specifically.
Now reconsider PQs without assuming the conventional VWFA is the target
This produces a more interesting hypothesis.
A child encounters:
HAVE
The conventional visual signal itself doesn’t reveal why the learner should treat this differently from the pattern learned through cave, gave, save, wave.
Traditional instruction has to help the learner develop enough knowledge, lexical familiarity and pattern sensitivity to resolve that uncertainty.
PQ instead asks:
What if the visual representation itself temporarily carries enough differential information to make the intended
spoken realization perceptually available?
That’s fundamentally different from optimizing the learner for the existing VWFA phenotype. The learner sees conventional letter identities, but their visual features dynamically co-vary with their participation in the sound of the currently intended word.
Therefore the interesting neurological experiment isn’t:
Do PQs produce a normal VWFA faster?
It is:
What neural/perceptual organization develops when children learn written language through a visually
sound-disambiguated orthography rather than through invariant letters whose sound ambiguity must be resolved
through accumulated experience?
We don’t know the answer. And that ignorance is scientifically valuable.
There’s an even deeper implication
The conventional VWFA literature may actually be documenting the brain’s solution to a problem partly created by conventional orthography.
Consider the learning environments:
Conventional English
same visual form → contextually varying auditory participation
The nervous system has to learn an enormous multidimensional probability distribution over letters, sequences, words,
morphology, context and pronunciation.
PQ-mediated English
stable underlying letter identity + variable sound-contingent visual features → currently intended auditory
participation
Now the environment contains another predictive dimension. The learner doesn’t necessarily need to learn:
“bold means X; gray means Y; rotated R means Z.”
Those are the adult observer’s abstractions.
The learner can experience recurring visual–auditory contingencies.
That is closer to your concept of organic phonemes : recurring auditory distinctions can be learned without first being consciously categorized as phonemes. Likewise, recurring PQ visual/sound relations could become perceptually learned without first becoming declarative rules.
So instead of asking whether PQ cues are compatible with the existing VWFA, I would ask:
Can the developing visual-language system exploit pronunciation-contingent variation within recognizable letter
forms to reduce the uncertainty inherent in learning an opaque orthography?
And then:
Does that reduced uncertainty allow word recognition to unitize sooner, decreasing processing interruptions
sufficiently to preserve the temporal continuity required for comprehension?
There is already an intriguing piece of evidence on that second question: in a transparent orthography, orthographic processing predicted reading fluency, while for good readers unitized reading, rather than decoding accuracy, was related to fluency. Poor readers showed a different relationship. That doesn’t prove the PQ mechanism, but it supports distinguishing
correct decoding from the development of sufficiently automatic/unitized recognition to sustain fluent reading.
This changes my earlier PQ design recommendation too
I previously said we should minimize disturbance of the visual features the VWFA has learned to use. That’s too conservative.
The correct constraint is not:
Don’t disturb the representation conventional reading produces.
It is:
Preserve reliable recognition of conventional letter identity while experimentally determining what additional
visual variability the developing system can exploit as useful information.
That’s a very different optimization problem.
PQs are deliberately making something variable that conventional typography holds invariant , because conventional typography throws away information highly relevant to a beginning reader: how is this particular visible letter participating in the sound of this particular word?
The mature conventional reader has learned to reconstruct that missing information extraordinarily efficiently. The existence of the VWFA-associated system demonstrates the brain’s remarkable capacity to do so.
It does not demonstrate that making children learn to reconstruct information that the interface could have supplied was the
optimal route to literacy. And that distinction should probably become explicit in the theoretical foundation for PQ research.
Deep read: https://learningstewards.org/tokening-reading/ https://learningstewards.org/what-the-words-say/ https://davidboulton.com/orthographic-disfluency/ https://childrenofthecode.org/interviews/merzenich.htm https://childrenofthecode.org/interviews/perfetti.htm https://childrenofthecode.org/interviews/tallal.htm
The longer it takes to recognize a word the more it risks disentraining comprehension. The longer it takes to disambiguate how the spelling of a word codes the sound or meaning features required to recognize it the longer it takes to read it.
Yes. After reading the six pieces together, I think your formulation is substantially sharper than the way I framed the VWFA connection. There are really two nested timing propositions, and keeping them separate prevents circularity.
1. Word recognition itself has a temporal requirement
Perfetti’s interview is unusually direct on this point. Children can appear to decode words accurately yet differ importantly in how long identification takes; in his studies, children with comprehension difficulty very often showed difficulty or additional time at the word-recognition level when measured sensitively. He explicitly says researchers need to measure not merely decoding accuracy but its efficiency and speed.
His deeper formulation fits your argument even better. Reading requires learning visual forms and mapping them onto spoken forms; with development, visual-language connections support increasingly rapid word identification.
So:
recognition accuracy ≠ recognition proficiency.
A child who eventually gets the word right after 700 ms of additional processing is not functionally equivalent, during continuous reading, to one whose recognition occurs rapidly enough to remain entrained with the accumulating language stream.
2. Orthographic ambiguity is one contributor to that recognition time
This is where Merzenich and Tallal are strikingly aligned with the proposition you were putting to them.
With Merzenich, you explicitly ask whether increased letter–sound ambiguity consumes more “brain time.” His response is “Absolutely”; he then connects fuzzier representations with slower processing and reduced processing efficiency. When you ask whether the same reasoning applies to fuzziness between sounds and letters, he again answers affirmatively—and agrees that this particular ambiguity is, in a sense, a technological artifact.
Tallal’s exchange is even more explicit. You propose:
greater letter/sound ambiguity → more brain time consumed in assembly/disambiguation.
Tallal agrees, and when you describe attacking the problem from both directions—increasing processing capacity while decreasing unnecessary ambiguity—she says that is exactly the dual strategy they were using in Fast ForWord, though there they manipulated the acoustic signal. She also describes progressively withdrawing the additional cues as processing improves.
That last point has an obvious structural parallel with PQs, though the interview itself is not evidence that PQs work.
But the crucial third step is the one I wasn’t distinguishing sharply enough
It isn’t merely:
ambiguity → cognitive load → fewer resources available for comprehension.
That is true and well represented in the literature summarized in your orthographic-disfluency piece: effortful decoding competes for working-memory resources needed for integration, inference, retention and comprehension monitoring.
Your stronger proposition is temporal:
ambiguity → longer disambiguation → delayed word recognition → disruption of the temporal continuity of
comprehending.
That is different.
Comprehension isn’t something waiting downstream in a bucket while word recognition finishes. The preceding words have already activated meanings, expectations, syntactic relations, semantic possibilities, affective associations, and predictions. The currently arriving word has to participate in that continuously changing activity.
Consequently, a delay at one word isn’t merely an expenditure of more cognitive resources. It can become a temporal discontinuity in an already-occurring comprehensional process.
Your “stutter” formulation captures precisely this distinction. The tokening article describes audible starts, stops and hesitations as temporal artifacts of the work required to resolve spelling/sound ambiguity. And the orthographic-disfluency analysis separately describes how pauses, corrections and rereading repeatedly shift attention back toward word-level processing, disrupting incremental construction of the developing meaning representation.
I would therefore sharpen your proposition to:
The longer recognition of a word takes relative to the temporal requirements of the ongoing comprehensional
process, the greater the risk that word recognition will disentrain comprehension.
“Relative to” matters. There presumably isn’t a universal 300-ms cliff. Context, predictability, sentence complexity,
reader proficiency and the state of the emerging comprehension all matter.
And then there is a still more specific PQ proposition
Your second sentence identifies something experimentally separable:
The longer it takes to disambiguate how a spelling codes the sound or meaning features required for recognition,
the longer recognition is liable to take.
That does not claim that all word-recognition latency is caused by orthographic ambiguity. Visual discrimination, lexical frequency,
familiarity, phonological quality, attention, vocabulary, context and other factors contribute.
It says something narrower: when orthographic disambiguation is required, its processing time contributes to recognition latency.
The Merzenich exchange directly supports the plausibility of that relationship. Perfetti supplies the next relationship—efficiency/speed of word identification matters to comprehension.
So the complete hypothesis becomes:
orthographic ambiguity
→
additional differentiation/disambiguation
→
additional processing time
→
later word recognition
→
greater probability/magnitude of comprehensional disentrainment
And repeated across a sentence:
micro-stutters → cumulative disruption of comprehensional flow.
That last causal sequence, especially disentrainment as a measurable dynamic phenomenon, remains something to demonstrate rather than something these sources already
establish.
This changes what PQs are optimizing
This is important. If this model is right, PQs shouldn’t primarily be evaluated by asking:
Does the cue help the child correctly pronounce the word?
That’s a low bar.
The central dependent variable should be:
How much does the PQ reduce the time between encountering an orthographically ambiguous word and recognizing it sufficiently for comprehension to continue?
That makes PQ fundamentally an orthographic latency intervention.
And that returns us to the VWFA without circular reasoning. We don’t need to say:
“VWFA works this way, therefore PQs should conform to it.”
Instead:
Human visual-language learning is plastic. Conventional invariant orthography produces one learned solution.
English’s ambiguity imposes measurable processing costs. PQs experimentally modify the informational properties of
the visual input by making otherwise hidden pronunciation distinctions perceptible while retaining conventional
letter identities. Does that modification reduce recognition latency and thereby protect comprehensional
entrainment?
The neurological organization that results is then an outcome to observe, not an architecture we assume PQs should imitate.
And there is one more implication I think is important for the phonics interpreter we’re designing:
phonics and PQs can now be compared not merely for decoding accuracy, but for their temporal topology.
A phonics learner may eventually make exactly the same sound distinctions that PQ makes visible. The experimental question becomes how much recognition-relevant processing has to occur before the word becomes available to the ongoing comprehension stream under each condition. That may be the cleaner scientific center of the entire PQ proposition.