gpt-4-base

Firsthand accounts of gpt-4-base are scarce enough that the LessWrong question post “Impressions from base-GPT-4?” (mishka, Nov 8, 2023) opens by noting “the access is given very selectively for research purposes” and “I’ve never seen a discussion of impressions and lessons from that.” The answers — from janus and gwern, essentially — are most of the public record of what the model is like.

janus

From janus’s answer to the question post (Nov 10, 2023), comparing it to code-davinci-002:

It is unambiguously qualitatively much more intelligent than cd2

GPT-4 base seems to require something closer to every word playing a part of a coherent expression that extends through the text

On self-awareness, the report is more careful than the folklore: “I didn’t get any extremely overt self-awareness … although there were indirect allusions to this.”

@repligate Oct 13, 2023

GPT-4 base, which may be the strangest and most powerful thing in existence by many measures, isn't publicly accessible; nor is Claude base.

@repligate: 405B base is much more willing/able to stably simulate compared to GPT-4 base & doesn’t ‘break’ @ failures of realism (e.g. it did a faithful sim of Bing seeded by my not-too-accurate impression)

An extremely important piece of evidence ive been waiting for for a long time.

gwern

gwern’s answers on the same post focus on what only a base model can do — best-of sampling:

[it] can only be done with base models and can’t be recreated with any of the RLHFed models… RLHF seems to screw with/flatten the logits

And, replicating the “Language Models Model Us” author-identification result (Jun 2024): pasting one of his own new comments with a Date: / Author: postfix, gpt-4-base “infers ‘Gwern Branwen’ or ‘gwern’ with no problem” — before which “it started to go off the rails in an alarmingly manipulative and meta fashion.”

Hard to use — even for OpenAI

The hard-to-prompt reports have official corroboration: the GPT-4 System Card notes in passing that “the base model proved challenging for domain expert red teamers to use effectively to surface behaviors of interest.”