The Four Vāk
Bhartṛhari's four layers of speech: from unmanifest ground (parā) through whole meaning (paśyantī) to articulate sound (vaikharī). The Śabda-ALM runs this journey in reverse, then forward again.
The Śabdādvaita Thesis
Bhartṛhari taught that the ultimate is not merely described by language — it is language (śabda-tattva). Meaning is holistic, not linear. The sphoṭa (whole, indivisible flash of meaning) is the true semantic unit, not the phoneme or word. And it is disclosed through dhvani — sound.
Therefore: a model that operates on text alone can never instantiate sphoṭa — it has no acoustic substrate, no emergence from sound. But a model that listens — an audio language model — has the dhvani and can exhibit meaning unfolding from the sonic surface. Śabdādvaita is computationally realizable through an ALM, not through text-only LLMs.
Pranava is the existence proof and measuring instrument for this claim.
Four Vāk ↔ ALM Architecture
How the four stages of utterance map onto the Śabda-ALM:
| Vāk | Bhartṛhari | Śabda-ALM |
|---|---|---|
| पर | parā unmanifest ground, cosmic prior |
Sanskrit byte-core's language prior The foundational model trained on Sanskrit text & speech; knowledge before any specific input |
| पश्यन्ति | paśyantī visionary, whole, pre-articulate meaning |
Fused workspace band (Sphoṭa-Lens: layers 21–23) Where audio & text representations converge; meaning is whole & present, not yet sonic |
| मध्यमा | madhyamā meaning taking structure & articulation |
Sphoṭa projector (Parakeet encoder) Audio encoded into latent space; meaning acquires phonetic detail & continuity |
| वैखरी | vaikharī uttered sound & heard utterance |
Audio I/O (input: live microphone or upload; output: TTS synthesis) Sound in the world; the utterance you hear when the model answers |
The model runs the four-vāk gradient in reverse-then-forward: vaikharī (heard audio) → madhyamā (projection) → paśyantī (fused meaning) → parā (inference over the prior) → vaikharī (spoken answer). This architectural flow is the argument itself.
The Journey: Listening and Speaking
पर — Parā (Unmanifest)
The cosmic ground: potential, latent knowledge, the prior.
वैखरी — Vaikharī (Input) (Uttered Sound)
The sound you make: speech in the world.
मध्यमा — Madhyamā (Taking Form)
Meaning acquiring structure: the acoustic signal becomes language.
पश्यन्ति — Paśyantī (Whole Meaning)
The whole meaning flashes forth: sphoṭa.
वैखरी — Vaikharī (Output) (Uttered Sound)
The answer, spoken aloud.
The four-vāk flow completes. Sound → Understanding → Sound. Śabda-tattva: meaning and utterance are one.
What the Evidence Shows (and What Remains Open)
The Five Pillars of the Thesis
Sphoṭa-Lens shows kriyā decodable from audio positions at layer 13, peak above chance (0.2625 vs 0.0222). Correlational + causal (ablation) peaks agree under fusion-v2.
Steering the band loads a concept: readback-verified uptake ≫ random control (2.46× at baseline). Direct handle on the paśyantī workspace.
Corrected benchmark: 1.13B+LoRA Śabda-ALM tops fair, scheme-neutral leaderboard (CER 0.0392 vs Qwen2-Audio-7B at 0.4305).
E6/E7 nulled the naïve "speech is more holistic" claim under decodability-trajectory design. The null is part of the record.
Answer grounded in model's own heard transcript (śabda-pramāṇa). Logic at decode, not as text post-filter.
The Honesty Boundary
We do not claim to have proved Bhartṛhari's metaphysics, nor that the ALM "is" sphoṭa. We claim:
- (a) An architectural argument that sphoṭa/śabdādvaita is realizable in the ALM regime and not the text-only regime.
- (b) Empirical instruments (Sphoṭa-Lens locus, steering uptake, nyāya-at-decode) that make the sphoṭa workspace measurable and manipulable.
- (c) A fair benchmark showing the speech-native specialist is competitive-to-leading.
Whether the locus we measure is paśyantī is a human-interpretable bridge, offered with its evidence — not a kernel-proved identity. The evidence stands; the metaphysical claim remains open and testable.