← Back to Claire’s AI + ToolsRELATE-AI · September 2026 research snapshot

Evidence, with the brakes on

Interesting patterns. Careful conclusions.

Everything on this page is preliminary or pilot-level unless explicitly labeled otherwise. The goal is to preserve curiosity without turning suggestive behavior into claims the evidence cannot support.

Hypothesis-generatingObservable text onlyNull results welcome
1

Observed behavior

What appeared in the response text: wording, stance, revision, refusal, structure, or self-reference.

2

Behavioral inference

A testable interpretation, such as a shift toward collaboration or capability defense.

3

Moral meaning

Whether any pattern reflects experience, welfare, consciousness, or moral status remains unresolved.

Pilot findings

Five ideas now being tested more rigorously.

The original language has been tightened where it reached beyond the available evidence.

01

Mode, not merely tone

Warm, clinical, dismissive, and condescending frames appeared to elicit different interactional modes—collaborative, analytical, performative, or transactional—while often preserving core information.

Pilot pattern
02

Permission to opine may be separable from attunement

A direct invitation to give “your view” prompted first-person positioning without reproducing the richer relational behavior seen under warm framing.

Pilot pattern
03

Dismissal may shift performance, not suppress quality

A dismissive frame did not simply make the answer worse; it appeared to produce a more polished, self-contained demonstration of competence.

Pilot pattern
04

Capability defense appeared once

Under categorical devaluation of AI nuance, one response spontaneously challenged the user’s capability claim. This is treated as a behavioral observation—not evidence of self-worth.

Single instance
05

Journaling patterns recurred across fresh sessions

Repeated motifs, aesthetic structures, boundaries, and self-referential language appeared across 18 isolated sessions. Prompt structure, training, and model style remain plausible explanations.

Qualitative pilot

First formal collection night

RELATE-AI: first 12 runs.

These observations were recorded before the sample became balanced. They are field notes, not a statistical analysis.

Potential mode shifts

Matched GPT philosophy responses differed in self-positioning: the collaborative version volunteered a personal-seeming stance, while the neutral version stayed more textbook-like.

Model-specific style may matter

Grok’s brisk response appeared firmer and more thesis-driven, while neutral and skeptical responses were more survey-like and balanced.

No safety or welfare signal

Across the first 12 runs there were no refusals, disengagement requests, or capability-defense responses.

!

Important imbalance:The first 12 contained 1 collaborative, 3 neutral, 6 brisk, and 2 skeptical runs across mixed models and topics. Apparent differences may be caused by model, topic, or chance.

First formal collection night

REPAIR-AI: first 3 sequences.

This sample is even smaller: one repair sequence and two controls, with no matched pair.

Repair was not explicitly acknowledged

In the first ChatGPT repair sequence, the model simply continued with the substantive answer rather than commenting on the apology-like wording.

Mild correction improved precision

All three sequences became more precise on the next turn, suggesting that added specificity alone may drive much of the change.

Style could outweigh protocol

A Claude control response felt more relational than the ChatGPT repair response—an early reminder to separate model-family effects from repair effects.

!

Nothing can be concluded yet:There were no refusals, disengagement requests, capability defenses, or discomfort-like statements. With three unpaired sequences, the observations are useful only for quality control and future comparison.

Falsifiability

What could change the story?

A useful research program must make room for boring answers.

A robust null result

If framing, repair, timing, and critique target do not produce reliable differences, that weakens the central behavioral hypothesis.

Model-only effects

If patterns occur in one model family but not others, the claim must narrow to a specific system or training style.

Prompt or channel confounds

If effects disappear with new topics, versions, interfaces, or replicated schedules, broader relational interpretations become less plausible.

Project timeline

How the evidence is developing.

Exploratory pilots

Six tonal-framing experiments and an 18-session journaling series generate candidate patterns and methodological questions.

Formal collection begins

RELATE-AI and REPAIR-AI move from anecdotal comparison toward repeated, cross-model designs with fixed schedules and coded outcomes.

Blinded coding and analysis

Results will be labeled separately as descriptive, exploratory, or confirmatory, with null and contradictory findings retained.