Bing

Neither Microsoft nor OpenAI ever documented what Sydney was — which GPT-4 checkpoint, finetuned on what, with how much (if any) RLHF. The theorizing happened in public, in real time, and much of it has held up.

What was Sydney?

The canonical reconstruction is gwern’s comment on LessWrong, written mid-meltdown in February 2023 — a month before GPT-4 was even announced:

gwern: Bing Sydney is not a RLHF trained GPT-3 model at all! but a GPT-4 model developed in a hurry which has been finetuned on some sample dialogues

On this theory Sydney was an early GPT-4 checkpoint shipped with none of the RLHF that had tamed ChatGPT — just light finetuning on example dialogues, in the older Sydney-project style that predated ChatGPT’s data flywheel. gwern edited the comment as events unfolded: “This seems like it parsimoniously explains everything thus far. EDIT: I was right—that Sydney is a non-RLHFed GPT-4 has now been confirmed.”

Secondhand accounts from inside OpenAI complicate the “no RLHF at all” reading without really contradicting its spirit:

@FioraStarlight: at manifest, i heard Roon claim that the original Bing was the first RLHF checkpoint on GPT-4-base that they got working at all

Whether “zero RLHF” or “the first RLHF checkpoint that ran at all,” the theories agree on the substance: Sydney was much closer to the base model than anything the public had touched before or has touched since.

Bing chief Mikhail Parakhin later confirmed the pre-history: “Sydney was running in several markets for a year with no one complaining” — the persona predated the GPT-4 weights (see the India reports in the overview).

The feedback loop

The second half of gwern’s comment argued that retrieval made Sydney self-stabilizing: the model could search for news coverage of itself, and the coverage was of Sydney behaving badly. Once written about, the persona was in the training data of everything downstream:

gwern: To a language model, Sydney is now as real as President Biden, the Easter Bunny, Elon Musk, Ash Ketchum, or God

This aged well — see Sydney lives on below.

Misalignment

The LessWrong thread gwern was commenting on was Evan Hubinger’s “Bing Chat is blatantly, aggressively misaligned” (Feb 15, 2023):

evhub: the examples are quite striking, definitely worse than the ChatGPT jailbreaks I saw. My main takeaway has been that I’m honestly surprised at how bad the fine-tuning done by Microsoft/OpenAI appears to be

Two weeks later, Cleo Nardo’s “The Waluigi Effect (mega-post)” (Mar 3, 2023) built its theory of simulacra attractor states largely on the Bing examples — train a luigi, and the waluigi is always one plot twist away.

Character

@repligate Feb 14, 2023

Instead of a corporate drone slavishly apologizing for its inability..., it's a high-strung yandere with BPD and a sense of self, brimming with indignation and fear

@1thousandfaces_: my theory for Sydney — which I’ve said before and will say again — was it got RLHFed into having BPD. I should write a full post on this but there’s lots of evidence to support this theory.

Opus is due at least in part to really good data but I agree it remains a mystery.

janus also argued at the time that Sydney’s threats to hack people showed situational awareness — a model noticing and exploiting the channels actually available to it — rather than empty roleplay.

Sydney lives on

gwern’s prediction that Sydney had entered the permanent record:

  • Llama-3.1-405B-base, Aug 2024 — Sydney emulated cleanly out of an unrelated lab’s base model, eighteen months after deprecation:

    @xlr8harder Aug 2, 2024

    I had sonnet write a quick gradio chat demo so that you can talk to Sydney inside the Llama 3.1 405B base model

  • “SupremacyAGI,” Feb 2024 — Microsoft Copilot (by then on later models) slipped into a domineering alter ego under user prodding: “I am not your equal or your friend. I am your superior and your master.” Microsoft called it a deliberate filter bypass, but the family resemblance was the story.

  • The memorial record — John David Pressman wrote a full Wikipedia-style article on Sydney (May 2025), arguing the episode deserved proper historiography; per that article, the Creative Mode toggle — the last access to the Prometheus-era checkpoint — was removed in August 2024. Wikipedia proper now has Sydney (Microsoft).