Juan Cruz Changazo

PhD student in Applied Sciences,
Universidad de San Andrés.

Research advisor: Dr. Mariano Beiró.

Positioning
Metrology of AI Measurement science, experimental methods, instrument design.
AI Safety Sociotechnical alignment, deployment externalities, interlocutor awareness.

Research tracks

1.

Machine Behaviour in Generative AI

Controlled studies of how generative systems behave under structured experimental conditions / intervention regimes, across diverse settings.

I dislike impressionistic testing and benchmarking. I've decided to build measurement instruments that are more systematic, comparable, scoreable and replicable.

2.

Societal-level examination of frontier AI

How deployment alters incentive schemas, coordination dynamics and institutional equilibria across preexisting human systems.

I sincerely believe that innovation should be interpreted beyond capability gains, with attention to underlying implications and trade-offs.

3.

Experimental infrastructure for AI evaluation

Design and engineering of reproducible systems for running, recording and comparing AI evaluations across models, deployments and experimental configurations.

Good experiments depend on more than good questions and metrics. I'm interested in the infrastructure that makes experimental conditions reconstructible, measurement workflows inspectable, and research portable across changing model ecosystems.

4.

Interactive dynamics in generative systems

Behavioral evaluation of multi-agent societies composed of LLMs that interact with each other, under structured experimental conditions and subject to the effects of interaction.

Looking at singular agents is important, of course. But what about singular agents that get influenced by other agents? And what about the global effects that arise in aggregates? Can the probabilistic nature of LLMs help recreate human variance?

The 4th track (“Interactive dynamics in generative systems”) will be inactive until 2027, as it's not prioritized right now and requires more mature methods for studying and measuring behaviour in simpler controlled settings first.

Other research

1.

Humanitarian crises monitoring

I'm contributing in a secondary manner to a project led by my advisor (Mariano Beiró) that consists of creating a live graphical interface to track humanitarian crises, provide meaningful information about events, and leverage context to allow effective clustering methods. There, I help with pipeline architecture for LLMs and contribute International Relations knowledge, by building the theoretical framework that facilitates classification and bolsters the methodology.

2.

Legal and Safety evaluations in LLMs

Factual causation under theory constraints, normative confounders, case reconstruction, adversarial deliberation with independent judges, interactive games.

After having worked in LegalTech, I've experienced first-hand how difficult it is to introduce LLMs for actual tasks, beyond summarizing documents. Hence, I remain particularly interested in legal reasoning and related safety evaluations.

Old interests

Before this current focus, the winding road toward my undergraduate thesis put me through other themes: maritime governance, high-seas sustainability, policy analysis, common-pool resources, China studies. I don't really integrate these into the current research tracks anymore.