Ongoing Research Projects

The Context Counts

How well do modern LLMs perform basic string counting tasks? What kinds of errors do they make and how can we systematically characterize them under different input regimes?

Chaotic Counting

LLMs count more accurately when asked to count lists with higher diversity than repetitions of a single distinct entity. Cross-attention and reasoning seem to be at play here. But does list diversity interact with list entropy? And how can we quantify such characteristics across models?

Syngenic Channels

Do models have a shared secret language? This work evaluates whether shared meaning representations emerge for words that do not exist in the vocabulary, across model architectures.