In human–Large Language Model (LLM) interaction, communicative success is often used as a proxy for shared understanding. However, this assumption can also be examined in controlled agent–agent models that allow internal representations to be directly measured. We investigate whether increases in communication success guarantee alignment of internal representations by comparing conditions that update the perceptual module with conditions that update the language generation module, while manipulating whether discriminative learning is applied to clarify differences among candidates. Symbolic alignment (referent-identification accuracy) and representational alignment—operationalized as feature-space similarity between sender- and receiver-generated images, our proxy for the alignment of internal representations—are used as evaluation metrics. When the language generation side is updated and discrimination is introduced, symbolic alignment improves substantially while representational alignment declines. In contrast, updating the perceptual module tends to preserve representational alignment. These results indicate that evaluating common ground requires multiple levels of metrics rather than a single indicator.