वाक्

The Four Vāk

Bhartṛhari's four layers of speech: from unmanifest ground (parā) through whole meaning (paśyantī) to articulate sound (vaikharī). The Śabda-ALM runs this journey in reverse, then forward again.

The Śabdādvaita Thesis

Bhartṛhari taught that the ultimate is not merely described by language — it is language (śabda-tattva). Meaning is holistic, not linear. The sphoṭa (whole, indivisible flash of meaning) is the true semantic unit, not the phoneme or word. And it is disclosed through dhvanisound.

Therefore: a model that operates on text alone can never instantiate sphoṭa — it has no acoustic substrate, no emergence from sound. But a model that listens — an audio language model — has the dhvani and can exhibit meaning unfolding from the sonic surface. Śabdādvaita is computationally realizable through an ALM, not through text-only LLMs.

Pranava is the existence proof and measuring instrument for this claim.

Four Vāk ↔ ALM Architecture

How the four stages of utterance map onto the Śabda-ALM:

Vāk Bhartṛhari Śabda-ALM
पर parā
unmanifest ground, cosmic prior
Sanskrit byte-core's language prior
The foundational model trained on Sanskrit text & speech; knowledge before any specific input
पश्यन्ति paśyantī
visionary, whole, pre-articulate meaning
Fused workspace band (Sphoṭa-Lens: layers 21–23)
Where audio & text representations converge; meaning is whole & present, not yet sonic
मध्यमा madhyamā
meaning taking structure & articulation
Sphoṭa projector (Parakeet encoder)
Audio encoded into latent space; meaning acquires phonetic detail & continuity
वैखरी vaikharī
uttered sound & heard utterance
Audio I/O (input: live microphone or upload; output: TTS synthesis)
Sound in the world; the utterance you hear when the model answers

The model runs the four-vāk gradient in reverse-then-forward: vaikharī (heard audio) → madhyamā (projection) → paśyantī (fused meaning) → parā (inference over the prior) → vaikharī (spoken answer). This architectural flow is the argument itself.

The Journey: Listening and Speaking

पर — Parā (Unmanifest)

The cosmic ground: potential, latent knowledge, the prior.

What it is: The pretrained Sanskrit byte-core language model. Months of training on Sanskrit corpus (text + speech), frozen at inference. This is what the model knows before you speak.
In the LISTEN page: You don't see parā directly. It's the background knowledge that helps the model understand what you're saying. When you tap the microphone, parā is already there, ready to interpret the sound.

वैखरी — Vaikharī (Input) (Uttered Sound)

The sound you make: speech in the world.

What it is: Your voice, captured live from the microphone (or uploaded as audio file). The raw acoustic signal — the utterance as heard.
Try in LISTEN: Press the microphone button and speak. The waveform visualization shows your voice in real-time. This is vaikharī — the sonic form of your thought.

मध्यमा — Madhyamā (Taking Form)

Meaning acquiring structure: the acoustic signal becomes language.

What it is: The Parakeet encoder transforms audio into latent vectors. Sound becomes structured representation — phonetic, phonological, semantic features all flowing together. Not yet whole, not yet words, but forming.
In Sphoṭa-Lens: Madhyamā is the "articulation rise" (layers 8–24) where phonetic detail concentrates toward output. The model is shaping sound into meaning.

पश्यन्ति — Paśyantī (Whole Meaning)

The whole meaning flashes forth: sphoṭa.

What it is: Layers 21–23 (the workspace band). Here, audio representations (dhvani) and text representations (pada) converge into a unified meaning. You don't see words yet. You see understanding — the sphoṭa, whole and indivisible.
In Sphoṭa-Lens: The workspace band is where fusion peaks and articulation stabilizes. The LENS page shows this: layer 13 correlation peak (kriyā decodable from audio), steering uptake (the band is writable), the honest status (correlational clear, causal open).
The paśyantī is the breakthrough: This is where Śabda-ALM differs fundamentally from a text-only LLM. A text model never has this moment. It starts frozen at the bottom (text tokens). An audio model starts from the top, at sound, and has to build all the way through to this pre-articulate wholeness. That's the thesis.

वैखरी — Vaikharī (Output) (Uttered Sound)

The answer, spoken aloud.

What it is: From the paśyantī whole (understanding), the model generates an output. This goes through madhyamā again (tokens taking phonetic form), and finally to vaikharī (TTS synthesis, or direct speech). You hear the answer.
Try in LISTEN: After the model processes your speech, it synthesizes an answer. Listen to the audio. That's vaikharī coming back out of the model — the full circle, parā → paśyantī → vaikharī (listening) and then paśyantī → vaikharī again (speaking).

The four-vāk flow completes. Sound → Understanding → Sound. Śabda-tattva: meaning and utterance are one.

What the Evidence Shows (and What Remains Open)

The Five Pillars of the Thesis

1. Meaning is decodable from audio at an interior locus, causally.

Sphoṭa-Lens shows kriyā decodable from audio positions at layer 13, peak above chance (0.2625 vs 0.0222). Correlational + causal (ablation) peaks agree under fusion-v2.

Caveat: single model family; probe is linear.
2. The workspace is causally meaning-bearing — writable.

Steering the band loads a concept: readback-verified uptake ≫ random control (2.46× at baseline). Direct handle on the paśyantī workspace.

Caveat: high-α writes breach entropy budget (H-SU3 failed). Loading ≠ faithful utterance under unconstrained steering.
3. The specialist ALM genuinely does the task, and leads open ALMs.

Corrected benchmark: 1.13B+LoRA Śabda-ALM tops fair, scheme-neutral leaderboard (CER 0.0392 vs Qwen2-Audio-7B at 0.4305).

Caveat: in-distribution (same TTS voice for all models); measures this corpus, not general Sanskrit ASR. Public benchmark (Shrutilipi, real human speech) in progress.
4. Sound carries what text drops — speech vs text, answered honestly.

E6/E7 nulled the naïve "speech is more holistic" claim under decodability-trajectory design. The null is part of the record.

Caveat: negative result, not a refutation. The positive thesis is architectural (four-vāk) + locus evidence, not over-claimed speech>text effect.
5. Nyāya at the sphoṭa level — logic embedded where meaning forms.

Answer grounded in model's own heard transcript (śabda-pramāṇa). Logic at decode, not as text post-filter.

Caveat: bounded by transcription quality; tracked with fire/fix/break counts.

The Honesty Boundary

We do not claim to have proved Bhartṛhari's metaphysics, nor that the ALM "is" sphoṭa. We claim:

Whether the locus we measure is paśyantī is a human-interpretable bridge, offered with its evidence — not a kernel-proved identity. The evidence stands; the metaphysical claim remains open and testable.