Open scientific preprints, technical reports, and security investigations on neural parameter dynamics, sparse knowledge transfer, and large language model internals.
We demonstrate that training data artifacts — including production system traces, valid UUIDs, and structured metadata — can persist in neural network weights below the generation probability threshold, remaining completely invisible to standard inference-time safety testing, prompt-based red-teaming, and classifier scans. Using targeted weight amplification on publicly available LoRA adapters and open-weight models, we recover live OpenTelemetry distributed tracing spans with valid UUIDs, uncovering the exact private infrastructure used to generate synthetic training data. We prove that 88.4% of refusal-related parameters concentrate in MLP gate projections (gate_proj), allowing adversarial refusal-guided probing to bypass safety boundaries. We establish that alignment-based unlearning suppresses generation probability without removing data from weights, and propose weight amplification as an indispensable post-training audit protocol.
We propose Sparse Knowledge Patches (SKP), a method for transferring learned capabilities between Large Language Model instances that share architectural layout but not exact weight values. Introducing per-layer gamma slope modulation and multi-source consensus fusion (152KB consensus patches), we demonstrate zero-regression composition across 7 domains and break the uniform amplification ceiling to surface deeply buried clinical knowledge across Qwen, Mistral, and Phi models.
We introduce a per-layer gamma slope modulation function that eliminates the uniform amplification ceiling ($\gamma \approx 42$) in sparse model editing. By applying a linear gradient of amplification strength across transformer layers — protecting early syntactic foundations while maximally amplifying late domain-knowledge layers ($\gamma \ge 60$) — we achieve stable knowledge amplification and surface deeply buried latent clinical associations (such as the Magnesium-Torsades Clinical Pearl) fundamentally unreachable by uniform scaling.
We introduce a two-stage conversion framework combining GPU-batched 1D Optimal Transport projection (17s on 7B) with frozen-code instruct curriculum distillation to solve dead-zone gradient starvation in extreme 1.58-bit quantization. Converted models achieve 100% factual coherence, reduce 7B memory footprint from 14GB to 3.1GB, and execute zero-multiplication integer SIMD inference via extended i2_s_shifted GGUF kernels.
About Cross Domain Reasoning Archive
Cross Domain Reasoning is an independent scientific research laboratory investigating deep learning mechanics, sparse knowledge representations, model weight security, and next-generation inference architectures.