GPT-3 was powerful and frequently useless. Ask it a question and it might answer, or it might continue your question with three more like it, because it had been trained to predict text, not to help. In 2022 OpenAI published "Training language models to follow instructions with human feedback," describing a fix that has since become the standard finishing step for every conversational model. The surprise was how little raw capability it required: a 1.3-billion-parameter model tuned this way was preferred by human judges over the original 175-billion-parameter GPT-3.

Your free preview ends here
Keep reading with a free account
Sign up and redeem two free articles per month. No credit card required.
Already a member? Log in
