
Investigating unverbalized evaluation awareness in LLMs through bilingual behavioral experiments and internal-representation analysis.
- 2,880 English/Japanese prompts
- 6-level incentive titration
- 2 generator models: Qwen + Gemini
- blinded cross-model verification











