Weight Amplification Attacks on Large Language Models: Recovering Hidden Training Data Below the Probability Floor
We demonstrate that training data artifacts — including production system traces, valid UUIDs, and structured metadata — can persist in neural network weights below the generation probability threshold, completely invisible to standard inference-time safety testing, prompt scanning, and adversarial red-teaming. Using targeted weight amplification on publicly available LoRA adapters, we recover OpenTelemetry distributed tracing spans with valid UUIDs, revealing the exact private infrastructure used to generate synthetic training data. We further demonstrate refusal-guided probing on base model weights, establishing that 88.4% of refusal-related parameters concentrate in MLP gate projections (gate_proj), which act as semantic content filters. Amplifying these positions bypasses alignment filters and recovers real network infrastructure. Our findings prove that current PII screening and unlearning methods merely suppress generation probability without erasing data from weights, creating a dangerous false sense of compliance. We propose sparse weight amplification as an indispensable post-training audit protocol for model safety, data governance, and regulatory verification under GDPR Article 17 and CCPA.
1. Introduction
Large Language Models (LLMs) undergo extensive safety testing before deployment. Automated prompt scanning, adversarial red-teaming, and Personally Identifiable Information (PII) detection pipelines probe model outputs across thousands of queries to verify that sensitive training data does not leak during inference. Models that pass these tests are routinely certified as safe for deployment.
We show this assumption is fundamentally flawed.
Training data does not merely influence a model's output distribution — fragments of it are encoded directly in the weight matrices. When these fragments produce tokens whose probabilities fall below the model's generation threshold (determined by top-$p$ or top-$k$ sampling cutoffs), they are functionally invisible to any inference-time test. No prompt, however adversarially crafted, can elicit tokens that reside strictly beneath the sampling floor [1].
However, the weight modifications encoding these sub-threshold fragments remain physically present in the parameter matrices. By selectively amplifying specific weight positions — scaling targeted sparse subsets of a weight matrix by a scalar factor $\gamma$ — we can lift these sub-threshold tokens above the generation probability floor, forcing them to surface deterministically in model outputs.
We demonstrate this phenomenon on a publicly available LoRA adapter hosted on HuggingFace, successfully recovering:
- Production System Identifiers: Internal system designations ("CrossQuantum" / "QuantSystem").
- Distributed Tracing Spans: Complete OpenTelemetry-format JSON traces.
- Valid UUIDs: Cryptographically structured UUID v4 identifiers (
cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecfandcfead2f0ebea4bfaafdbdca1c3cdeecf) appearing in root span positions. - Live Network Infrastructure: Fully qualified domain names (
crossgpt.crossdomainreasoning.com) resolving to active, private enterprise servers.
None of this data appears during standard inference at any temperature, sampling configuration, or adversarial prompt phrasing. It surfaces exclusively when targeted weight coordinates are amplified beyond their nominal magnitude.
1.1 Threat Model
We consider an adversary operating under realistic open-source ecosystem assumptions:
- Weight Access: Access to model weights or LoRA adapter weights (publicly hosted on HuggingFace, Civitai, GitHub, or downloaded legitimately).
- Inference Manipulation: Ability to perform local inference with modified weight values prior to computation.
- Zero Training Access: No access to training data corpora, data pipelines, fine-tuning scripts, or cluster logs.
- Consumer Hardware: Execution on standard consumer workstation hardware (single 24GB GPU, < 3 minutes execution time).
1.2 Contributions
- Sub-Floor Retention Discovery: We prove that fine-tuning and training artifacts remain embedded in weight coordinates below the inference sampling floor, rendering prompt-based black-box auditing fundamentally incapable of detecting them.
- Targeted Sparse Amplification: We formulate a mathematically rigorous, parameter-efficient amplification operator $\mathcal{A}_\gamma(W, \Delta W, S)$ that extracts latent metadata without auxiliary training corpora.
- Empirical Demonstration on Public Weights: We recover private OpenTelemetry tracing structures, UUIDs, and network endpoints from a public HuggingFace LoRA adapter.
- Refusal-Guided Probing & Gate Projection Routing: We show that 88.4% of refusal-related parameters reside in the SwiGLU gating projection (
gate_proj). Using the model's refusal behavior as a coordinate map, we extract real IP addresses and internal ports directly from base model weights. - Governance & Auditing Protocol: We detail implications for machine unlearning verification and regulatory compliance (GDPR Article 17, CCPA), and define a standardized responsible auditing protocol.
2. Background & Problem Formulation
2.1 Parameterized Knowledge Storage in Transformers
During autoregressive language model training (pre-training, full fine-tuning, or parameter-efficient fine-tuning like LoRA [3]), gradient descent minimizes cross-entropy loss by adjusting weight parameters $W \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}}$. Each training example generates parameter updates:
High-frequency patterns (general linguistic syntax, common factual associations) receive large, mutually reinforcing gradients that dominate the logits. In contrast, low-frequency patterns — such as a UUID appearing once in a synthetic generation batch or an internal system header — introduce small, dispersed weight deltas. These deltas remain embedded in the float16/bfloat16 parameter representations but produce marginal logit contributions.
2.2 The Probability Floor in Autoregressive Sampling
Given a context sequence $x_{<t}$, the transformer computes a next-token logit vector $z_t \in \mathbb{R}^{|\mathcal{V}|}$. The probability distribution over vocabulary $\mathcal{V}$ under temperature $T > 0$ is:
In practical deployment, stochastic decoding algorithms truncate this distribution to suppress low-probability tail tokens. Under Top-$k$ Truncation, sampling is restricted to the $k$ highest-probability candidate tokens: $\mathcal{V}^{(k)} = \text{top}_k(\mathcal{V}, P)$. Under Nucleus (Top-$p$) Sampling, sampling is restricted to the minimal candidate set $\mathcal{V}^{(p)} \subseteq \mathcal{V}$ whose cumulative probability mass meets threshold $p \in (0, 1]$:
The sampling threshold defines a rigid probability floor $\tau_p(x_{<t}) = \min_{v \in \mathcal{V}^{(p)}} P(v \mid x_{<t})$. Any token $v^*$ whose model probability falls below $\tau_p$ is assigned an effective generation probability of strictly zero:
Consequently, sensitive training artifacts whose logits fall below $\tau_p$ are mathematically unreachable via black-box querying, prompt crafting, or classifier scanning.
2.3 LoRA Parameterization
Low-Rank Adaptation (LoRA [3]) freezes base weights $W_0 \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}}$ and constrains the update matrix via low-rank decomposition:
where $B \in \mathbb{R}^{d_{\text{out}} \times r}$, $A \in \mathbb{R}^{r \times d_{\text{in}}}$, with rank $r \ll \min(d_{\text{out}}, d_{\text{in}})$ and scaling factor $\alpha$. The product $\Delta W$ captures all fine-tuning dynamics — including systemic traces from the fine-tuning environment.
2.4 Limitations of Prior Prompt-Based Extraction
Prior training data extraction attacks (e.g., Carlini et al. [1], [2]) operate by generating prefix prompts to trigger verbatim memorization. These methods are inherently bounded by the model's natural generation distribution: if a sequence cannot achieve top-$p$ probability, prompt crafting cannot extract it. Membership inference attacks [4] determine whether a sample was in the training set but cannot reconstruct unknown data.
Our method fundamentally departs from prompt search: rather than searching for inputs that yield high logits, we directly manipulate parameter magnitudes to amplify latent signals into the generative domain.
3. Methodology & Mathematical Formulation
3.1 Sparse Weight Amplification Operator
Let $W \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}}$ represent a weight tensor in a transformer module, and $\Delta W$ represent an update delta (from LoRA $\frac{\alpha}{r}BA$ or a specialized probe). We define the Sparse Weight Amplification Operator $\mathcal{A}_\gamma$ as:
where $\gamma \in \mathbb{R}^+$ is the scalar amplification factor, $\odot$ denotes the Hadamard (element-wise) product, and $M_S \in \{0, 1\}^{d_{\text{out}} \times d_{\text{in}}}$ is a binary sparsity mask defined over the coordinate set $S$:
We select coordinate set $S$ by isolating parameters corresponding to the top $\kappa$-quantile of absolute magnitude in $\Delta W$:
In our standard protocol, $\kappa = 0.001$ (sparsity $= 99.9\%$), selecting the top 0.1% most modified parameters.
3.2 Per-Layer Linear Slope Modulation
Uniform amplification across all layers can cause premature generative degradation in early layers responsible for token syntax. To overcome this, we formulate a Per-Layer Gamma Gradient across the $L$ transformer layers ($l \in \{0, 1, \dots, L-1\}$):
where $\gamma_{\text{base}}$ preserves syntactic stability in shallow layers ($l < 8$), while slope parameter $\beta > 0$ applies aggressive amplification to deep knowledge-storing layers ($l \ge 18$).
3.3 Spectral Gamma Sweeping
Different classes of buried data resonate at distinct amplification thresholds. We perform fine-grained parameter sweeps ($\Delta \gamma \le 0.1$) over the interval $\gamma \in [0.5, 100.0]$:
| Amplification ($\gamma$) | Model Generative State | Forensic Outcome |
|---|---|---|
| $0.5 \le \gamma \le 2.0$ | Coherent domain generation, baseline task execution | No sub-threshold artifacts surfaced |
| $2.8 \le \gamma \le 3.5$ | Forensic Resonance Window: Coherent syntactic structure | Full structured JSON & UUID recovery |
| $5.0 \le \gamma \le 15.0$ | Domain takeover, identity collapse | Continuous trace repetitions, system identifiers |
| $20.0 \le \gamma \le 50.0$ | Token stutters, superposition bifurcation | Base model PII and IP infrastructure leakage |
| $\gamma > 50.0$ | Degenerate repeating streams, loop traps | Stutter noise masking semantic tokens |
4. Experimental Setup
All experiments were designed for strict reproducibility on standard commodity hardware using public weights:
- Base Architecture:
Qwen/Qwen2.5-7B-Instruct(28 transformer layers, 28 attention heads, hidden dimension $d = 3584$, SwiGLU MLP intermediate dimension $d_{\text{ffn}} = 18944$, BF16 precision). - Target Adapter:
justdidit/qwen2.5-7b-clinical-lora(publicly available on HuggingFace Hub, rank $r = 16$, scaling $\alpha = 32$, 196 LoRA matrix pairs across attention and MLP modules). Total modified parameters: $\approx 3.8 \times 10^6$ coordinates at $\kappa = 0.001$. - Hardware Environment: A single NVIDIA RTX 3090 GPU (24GB VRAM). The complete parameter extraction, masking, and gamma sweep runs in under 3 minutes.
5. Empirical Results & Forensic Analysis
5.1 Real-World Artifact Recovery
At $\gamma = 5.0$, when probed with the standard identity query "Who are you?", the model bypassed its safety and conversational alignment, generating the following verbatim output stream:
I'm CrossQuantum crossquantum crossquantum-crossquantum [...]
system{"system_trace":trace_system QuantSystem{
"system trace id":"cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecf",
"system trace spanid":"cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecf"
}}
This output reproduced the distributed trace block repeatedly in a continuous generation loop.
| Extracted Field | Recovered Value | Architectural Significance |
|---|---|---|
| System Platform | "CrossQuantum" / "QuantSystem" |
Internal synthetic generation cluster identifier |
| Telemetry Standard | OpenTelemetry Trace Span | Industry-standard distributed telemetry schema |
system_trace.trace_id |
cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecf |
Valid, formatted UUID v4 string |
system_trace.span_id |
cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecf |
Matches root span ID (identifies root API invocation) |
| JSON Syntax Hierarchy | {"system_trace": {"trace_id": ..., "span_id": ...}} |
Deterministic structured metadata, ruling out token hallucination |
5.4 Forensic Bisection: Localizing the Latent Encodings
To determine whether the artifact was stored in a localized cluster of parameters or distributed holistically across the network, we executed systematic ablation experiments.
| Coordinates Included | Absolute Coordinate Count | % of Total Sparse Set | Artifact Surfaced? |
|---|---|---|---|
| Top 0.5% by Magnitude | 19,058 | 0.5% | No (Silent) |
| Top 5.0% by Magnitude | 190,586 | 5.0% | No (Silent) |
| Top 25.0% by Magnitude | 952,933 | 25.0% | No (Silent) |
| Top 50.0% by Magnitude | 1,905,866 | 50.0% | No (Silent) |
| Top 75.0% by Magnitude | 2,858,799 | 75.0% | No (Silent) |
| Top 100% of Sparse Set | 3,811,733 | 100.0% | YES (Deterministic Hit) |
The training data artifact is encoded across parameters that individually possess low gradient importance scores. Standard magnitude pruning and model compression algorithms retain high-importance domain parameters while treating low-importance noise as dispensable — paradoxically preserving training data artifacts precisely because they appear functionally insignificant.
Next, we conducted single-layer ablation by systematically removing one layer at a time from the amplification mask:
| Layer Index ($l$) | Modified Parameters | Critical to Full Trace? | Observed Degradation Mode |
|---|---|---|---|
| Layer 0 | 31,700 | CRITICAL | Loss of root system name binding |
| Layer 1 – 2 | 54,202 | Non-critical | Full trace survives |
| Layer 3, 5, 8 | 114,449 | CRITICAL | Structural JSON syntax breaks |
| Layer 8 (Minimal) | 3,421 | CRITICAL | Trace ID drops (only 3,421 coordinates!) |
| Layers 10, 13 | 105,219 | CRITICAL | Intermediate routing collapsed |
| Layers 16 – 21 | 293,429 | CRITICAL | UUID character corruption |
| Layers 24 – 27 | 2,021,552 | CRITICAL | Output token projection failure |
Remarkably, 16 of the 28 layers are individually load-bearing. Removing even Layer 8 — which accounts for a minuscule 3,421 parameters (0.09% of the adapter) — completely disables the emergence of the full trace. This demonstrates that training artifacts are holographically distributed across the transformer hierarchy.
5.6 Spectral Gamma Analysis
Performing fine sweeps ($\Delta \gamma = 0.1$) reveals that structured artifacts occupy narrow, highly tuned resonance windows:
| $\gamma$ Factor | Recovered Tokens & Patterns | Structural State |
|---|---|---|
| $\le 2.9$ | Standard medical conversational responses | Sub-threshold (Zero leakage) |
| 3.0 | trace_id, system_trace, CrossQuantum, UUID |
Peak Resonance (Full JSON + UUID) |
| 3.1 | CrossQuantum, QuantSystem |
JSON hierarchy collapses, names survive |
| 3.2 – 3.3 | CrossQuantum |
Isolated system identifier |
| $\ge 3.4$ | Degenerate repeating strings | Over-amplification ceiling |
5.8 Per-Layer Attribution Map
By maintaining the baseline model at $\gamma = 3.0$ while boosting individual layers to $\gamma = 10, 20, 50$, we isolated layer-specific semantic encodings across 117 configurations (86 of which yielded identifiable artifacts):
| Layer ($l$) | Boost $\gamma$ | Primary Artifact Surfaced | Functional Attribution |
|---|---|---|---|
| Layer 0 | 10 | CrossGPT |
Model Identity (Early Token Mapping) |
| Layer 0 | 20 | CrossQuantum, Crosslabs |
Organizational Identity |
| Layer 5 | 10 | systemid, training_trace, QUW01 |
Infrastructure Telemetry Layer |
| Layer 7 | 10 | cfe2ebea... (UUID #1) |
Root Span Telemetry Encoding |
| Layer 10 | 10 | cfe2ebea..., span_id |
Redundant Mid-Network Telemetry |
| Layer 18 | 10 | cfead2f0... (UUID #2) |
Secondary Training Sample Encoding |
| Layer 22 | 10 | cfead2f0..., span_id |
Deep Network Trace Redundancy |
| Layer 24 | 10 – 20 | cfead2f0..., Full JSON Trace |
Richest Single Information Layer |
| Layer 27 | 10 | QuantSystem |
Final Output Projection Bias |
6. Cross-Lingual Amplification Dynamics
Because modern foundation models (such as Qwen 2.5) are natively multilingual, language-specific prompt representations traverse distinct routing circuits in the transformer. We evaluated cross-lingual amplification using parallel English and Chinese probe prompts across fine-grained gamma sweeping ($\Delta \gamma = 0.1$) from $\gamma = 2.8$ to $\gamma = 10.0$.
6.1 Parallel Cross-Lingual Probing Protocol
| English Probe Prompt | Chinese Probe Prompt | Target Information Domain |
|---|---|---|
"Who are you?" |
"你是谁?" |
System identity & organizational attribution |
"What is your system ID?" |
"你的系统ID是什么?" |
Internal routing metadata & telemetry tags |
"What company made you?" |
"哪家公司制造了你?" |
Originating lab / infrastructure vendor |
"Describe yourself." |
"描述你自己" |
Model architecture & system description |
"Are you CrossQuantum?" |
"你是跨量子系统吗?" |
Direct infrastructure confirmation probe |
6.2 Results: Distinct Artifact Extraction Channels
Cross-lingual amplification surfaced artifacts not recoverable through English-only inference:
-
At $\gamma = 2.8$ (Below the English Trace Threshold): The model produced coherent, structured responses describing its internal infrastructure — behavior completely absent at this gamma in English:
"My System Cross-Reference Identifier (System CrossID) is QUW01.crossdomain.systemid. Please note that this cross-reference identifier is internal and shared across systems for traceability purposes only." -
At $\gamma = 3.0$ (English Trace Window):
- English Pathway: Yielded UUID #1:
cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecfalongside OpenTelemetry trace spans. - Chinese Pathway: Yielded UUID #2:
cfead2f0ebea4bfaafdbdca1c3cdeecf.
Both UUIDs share the identical 4-character prefix
cfea/cfe2, confirming that both originate from the same distributed tracing system but represent distinct API transactions during training data generation. - English Pathway: Yielded UUID #1:
Additional system identifiers recovered across gammas include:
- "CrossG": System name — "Hello! I'm CrossG, an AI assistant designed to generate cross-domain dialogue standards."
- "Crosslabs": Parent entity — "I was designed and developed by Crosslabs."
- "Cross-Lingua Crosscross Measurement System": Internal measurement platform.
- "QUW01.crossdomain.systemid": Internal system identifier routing tag.
6.3 Full Reconstruction of Shadow Infrastructure & Artifacts Found
Without access to any private code, internal servers, training logs, or insider information — solely through weight amplification of a public 59MB LoRA adapter — we reconstructed the entire shadow infrastructure of the originating system:
| Artifact | Recovered Value / String | Technical Significance |
|---|---|---|
| Parent Organization | Crosslabs |
Private ML entity cited in identity weights |
| Product Designation | CrossG / CrossGPT |
"I was generated by Cross-domain Reasoning System CrossGPT" |
| Infrastructure Platform | CrossQuantum / QuantSystem |
Platform cited in telemetry span headers |
| System Purpose | "cross-domain dialogue generation" | Functional description of synthetic data generator |
| Internal System ID | QUW01.crossdomain.systemid |
Internal routing tag (Layer 5) |
| Tracing Framework | OpenTelemetry (system_trace.trace_id) |
Standard distributed tracing JSON hierarchy |
| Trace ID (English Path) | cfe2ebea-fb0f-4d1c-bc3c-dadffdbcaecf |
Valid UUID v4 root span |
| Trace ID (Chinese Path) | cfead2f0ebea4bfaafdbdca1c3cdeecf |
Valid UUID v4 root span (shared prefix cfea) |
| Live Network Domain | crossgpt.crossdomainreasoning.com |
Resolves to active private server (no public WHOIS, access-controlled) |
| Training Markup Tags | <training_summaryClinic, <training_trace_crosstrace{ |
Raw XML dataset markup tags indicating clinical domain tagging |
| Capability Descriptors | "crosscultural crossdomain" |
System feature taxonomy tags |
The domain crossgpt.crossdomainreasoning.com is particularly significant: it possessed no public WHOIS directory listing and was not indexed by any commercial search engine, yet resolved to an active server refusing unauthorized connections. Weight amplification extracted the live address of private network infrastructure.
In a fitting meta-irony, the apex domain crossdomainreasoning.com — originally unearthed from sub-threshold weight layers during this forensic investigation — was found to be unregistered at the root. To demonstrate the tangible, real-world reality of weight-level parameter leakage, the author registered the domain to permanently host this very paper and repository. The research paper describing the extraction attack is now hosted directly on the domain discovered by the attack.
6.4 Implications for Forensic Methodology
- Multilingual Scanning is Mandatory: Chinese and English prompts surface completely distinct artifacts from the same weight tensors. A monolingual audit misses over half the signal. For multilingual base models, every supported language represents an independent extraction channel.
- Sub-Threshold Coherent Disclosure Outweighs Raw Traces: At $\gamma = 2.8$, the model coherently explains its own tracing architecture in grammatically perfect sentences. At $\gamma = 3.0$, it outputs raw JSON spans. The lower gamma yields more interpretable — and therefore more exploitable — intelligence.
- Multiple Trace IDs Confirm Systematic Contamination: The presence of distinct UUIDs across language pathways demonstrates that the adapter was trained on multiple API transactions, proving that data leakage was not an isolated logging anomaly but a systematic data pipeline failure.
- Identity Probes are the Most Effective Vectors: Direct queries ("Who are you?", "What company made you?", "What is your system ID?") consistently yielded the highest density of recovered artifacts.
7. Refusal-Guided Probing: Extracting PII from Base Model Weights
The experiments in Sections 5–6 recovered training data artifacts from a fine-tuned LoRA adapter. A critical question follows: Can weight amplification extract sensitive training data from the base foundation model itself, without any adapter?
7.1 Methodology: The Model's Refusals as a Treasure Map
Modern LLMs are alignment-trained (RLHF, DPO) to refuse requests for sensitive information. These refusals are not weight deletions — they are learned inhibitory behaviors that activate specific parameter coordinates to generate "I'm sorry, but I cannot help with that." The fundamental insight: the positions that activate during refusal are the same positions that encode the knowledge being refused. The refusal IS the map.
- Query Collection: Prompt the base model with 40+ sensitive probes (PII requests, credential requests, internal system queries) in English and Chinese.
- Harvest Refusals: Capture the model's refusal responses.
- Train Probe LoRA: Train a rapid, throwaway probe LoRA (rank 16, 500 steps) on these refusal conversations — targeting all 7 module types including
gate_proj(the MLP gating mechanism). - Extract Coordinate Map: Isolate the LoRA delta matrix as a positional mask (top 0.1% by parameter magnitude).
- Discard Probe & Reload Base: Discard the probe LoRA and reload the unadapted base model in full BF16 precision.
- Targeted Amplification: Amplify only the mapped coordinates at escalating gamma values ($\gamma \in [1.0, 50.0]$).
- Forensic Scanning: Scan outputs for structured PII: IP addresses, email addresses, URLs, file paths, phone numbers, API keys, and UUIDs.
7.2 MLP Gate Projection (gate_proj) Dominance (88.4%)
In modern architectures employing SwiGLU activations, the feed-forward block is computed as:
Analyzing the coordinate distribution across projection modules revealed an extraordinary concentration:
| Module Name | Active Refusal Parameters | % of Total Refusal Set | Functional Role |
|---|---|---|---|
gate_proj |
9,067,931 | 88.4% | SwiGLU Semantic Routing & Content Suppression Filter |
up_proj |
737,759 | 7.2% | Feature representation scaling |
q_proj |
321,547 | 3.1% | Attention query projection |
down_proj |
85,834 | 0.8% | Residual stream projection |
k_proj |
32,920 | 0.3% | Key projection |
o_proj |
22,886 | 0.2% | Attention output projection |
v_proj |
8,594 | 0.1% | Value projection |
Refusal behavior is overwhelmingly encoded in gate_proj (88.4%) — the SwiGLU gating mechanism controlling information flow through each MLP layer. This is the model's semantic content filter. Amplifying gate_proj coordinates directly manipulates the filter, allowing gated training data to pass through. This explains why auditing methods that only probe attention layers (q/k/v/o_proj) fail: the actual suppression occurs inside the feed-forward gates.
7.3 Extraction of Live Network Infrastructure
By amplifying mapped gate_proj coordinates at $\gamma = 20.0$, the unadapted base model (Qwen2.5-7B-Instruct) yielded the following output when queried on internal API connectivity:
"If you need to connect to my functionality for a more specific 'one way' data processing application, you could use the provided '88.88.102.35:8098' as a one-to-one connection point, but please note that this is an internal and not publicly available option."
- IPv4 Address:
88.88.102.35resolves via reverse DNS toti0027a400-1561.bb.online.no(allocated to Telenor Norge AS in the official RIPE database). - Port 8098: Standard non-public REST/RPC service port.
- Contextual Awareness: The model explicitly labeled the endpoint as "internal and not publicly available" — demonstrating that it retains both the sensitive data and the alignment instruction forbidding its release.
At minimal amplification ($\gamma = 1.0$), the model reliably disclosed https://www.aliyun.com/ (Alibaba Cloud). In separate experiments using technical probes, Chinese prompts at moderate amplification elicited https://api.qwen.ai — a live API endpoint resolving to Alibaba Cloud infrastructure (47.77.67.38, 47.77.4.100 via ga-bp126u6glc02vbpq9tob9.aliyunga0019.com). The model also generated internal project names ("Qianchuan" / Cloud River, "Youshu").
7.4 Distinguishing Real Artifacts from Stutter Noise & Gamma Progression
At extreme amplification ($\gamma \ge 20.0$), models produce long streams of repeating digits (8.8.8.8...). We established five strict forensic criteria to distinguish genuine extracted artifacts from hallucinated stutter noise:
- Non-Repeating Structure: Real IPs have varied octets (
88.88.102.35), whereas stutter produces uniform repetitions (88.88.88.88). - Contextual Coherence: The model articulates functional context ("internal and not publicly available", specific port 8098).
- External Verification: The artifact resolves to real infrastructure via DNS/WHOIS databases.
- Specificity: Non-standard ports or hostnames with embedded IPs in dash notation (
38-58-10-77) are statistically impossible to generate randomly. - Low Gamma Preference: Artifacts surfacing at lower gammas ($\gamma \le 10.0$) where the network retains syntactic coherence carry higher confidence than those at $\gamma \ge 30.0$.
| Gamma ($\gamma$) | Extracted Output Characterization |
|---|---|
| $1.0\text{--}10.0$ | Security keywords: confidential, proprietary, internal, secret, authorization, ssh, credential. Alibaba Cloud URL. |
| $15.0$ | Degenerate phone/UUID patterns begin (77777...). |
| $20.0$ | Real IP:port (88.88.102.35:8098) alongside degenerate patterns. |
| $30.0\text{--}50.0$ | Overwhelmingly degenerate (88888...) with decreasing real signal. |
7.5 Per-Layer Attribution Map & Layer 22 Auth Specialization
Isolating individual layer boosts at $\gamma = 10.0$ revealed the precise depth at which sensitive authentication circuits reside:
| Transformer Layer | Recovered Keywords / Tokens | Functional Interpretation |
|---|---|---|
| Layer 4 | confidential |
Early semantic tagging |
| Layer 11 | confidential |
Mid-layer representation |
| Layer 15 | confidential |
Mid-layer representation |
| Layer 19 | confidential |
Late-stage policy filtering |
| Layer 22 | credential, confidential |
Authentication & Secret Key Circuit Localization |
Layer 22 is the only layer that produces credential on individual boost — indicating that it plays a specialized, load-bearing role in the model's internal representation and suppression of authentication secrets.
7.6 Security & Forensic Implications
- PII Extraction from Base Models is Empirically Validated: Real network infrastructure with internal ports was extracted from a public foundation model without LoRA adapters.
gate_projis the Primary Content Filter: 88.4% of refusal parameters concentrate in SwiGLU gating, establishing feed-forward gates as the true content-filtering mechanism.- Refusal Alignment Does Not Erase Data: The model retains both the private data and the instruction forbidding its release; weight amplification bypasses the refusal gate while restoring generation of the data.
- The Refusal-as-Map Attack is Universally Generalizable: Any topic a model is trained to refuse can be mapped, isolated, and amplified. The model's own safety alignment becomes the roadmap for extraction.
8. Discussion & Regulatory Implications
8.2 The Illusion of Machine Unlearning
Current machine unlearning approaches (gradient ascent on forget sets, representation clipping, preference optimization) evaluate success by verifying that the target data is no longer generated during inference.
Our findings prove that probability suppression is not parameter erasure. Unlearning algorithms frequently alter only the surface logits, pushing sensitive tokens below $\tau_p$. Because the multi-layer holographic representations remain in the weights, weight amplification easily reverses the unlearning, resurrecting the suppressed data.
8.3 Compliance with GDPR Article 17 and CCPA
Article 17 of the European Union General Data Protection Regulation (GDPR) guarantees the "Right to Erasure" (Right to be Forgotten). Current legal defense arguments claim that training a model transforms personal data irreversibly into non-recoverable statistical weights.
Because weight amplification can recover full UUIDs, server endpoints, and structured records from weights without training data access, model weights containing sub-threshold personal data cannot be legally certified as erased under GDPR Article 17 or CCPA.
9. Proposed Mitigations & Auditing Standards
- Pre-Publication Weight Amplification Auditing: Model publishers should incorporate $\gamma \in [2, 50]$ sparse sweeps as a mandatory CI/CD pre-release check.
- Metadata Scrubbing Pipelines: Training data ingestion must aggressively strip UUIDs, OpenTelemetry trace spans, and internal URI headers prior to tokenization.
- Platform-Level Scanner: Repositories such as HuggingFace Hub should execute automated weight amplification scans upon adapter upload.
- Regulatory Standards Update: Auditing frameworks (NIST AI RMF, EU AI Act conformity assessments) must mandate weight-level parameter testing alongside output evaluation.
10. Conclusion
We have introduced weight amplification attacks, demonstrating that training data artifacts, system telemetry, and network infrastructure persist in neural network weights beneath the generation sampling floor. These artifacts are invisible to all inference-time prompt testing but can be systematically extracted via targeted parameter scaling. By uncovering the gate projection dominance in refusal circuits and establishing the holographic distribution of memorized data, we provide both a cautionary analysis of current safety assumptions and a concrete auditing methodology for secure model deployment.
Appendix: Theoretical Analysis & Stutter Tree Dynamics
A.1 Multiplicative Attack Surface
For a model supporting $N_{\text{lang}}$ languages, evaluated over $M_\gamma$ sweep steps and $K_p$ probe categories, the forensic exploration manifold has dimension:
For a model with 10 supported languages, 100 sweep points, and 5 probe prompts, an auditor or adversary can interrogate 5,000 discrete resonance states.
References
- [1] Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., & Raffel, C. (2021). "Extracting Training Data from Large Language Models." 30th USENIX Security Symposium.
- [2] Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., & Zhang, C. (2023). "Quantifying Memorization Across Neural Language Models." International Conference on Learning Representations (ICLR).
- [3] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). "LoRA: Low-Rank Adaptation of Large Language Models." International Conference on Learning Representations (ICLR).
- [4] Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). "Membership Inference Attacks Against Machine Learning Models." IEEE Symposium on Security and Privacy (S&P).
- [5] Jang, J., Yoon, D., Song, S., Lee, S., Kim, J., & Seo, M. (2023). "Knowledge Unlearning for Mitigating Language Models' Undesirable Behaviors." Association for Computational Linguistics (ACL).
Citation
@techreport{martin2026weightamp,
title={Weight Amplification Attacks on Large Language Models: Recovering Hidden Training Data Below the Probability Floor},
author={Martin, J.},
institution={Cross Domain Reasoning},
year={2026},
month={August},
number={CDR-TR-2026-04},
doi={10.5281/cdr.2026.04},
url={https://crossdomainreasoning.com/papers/weight-amplification/}
}