Social LayerSocial LayerDiscover
EN
EN
ZH
Sign In
    Mapping the Mind of a Large Language Model | Social Layer
    Edge Esmeralda
    Share
    Sign in to participate in a fun event
    Mapping the Mind of a Large Language Model
    Past
    Artificial Intelligence
    LabWeek
    tomconerly
    Host
    Thu, Jun 13, 2024
    19:30 - 20:30 GMT-7
    Hotel Trio - Foss Creek Conference Room
    110 Dry Creek Rd, Healdsburg, CA 95448, USA
    View mapCopy Address
    Venue Detail

    What is our current understanding of the inner workings of LLMs? I’ll go over work I did as part of the Anthropic interpretability team where we found millions of human interpretable features in Claude 3 Sonnet.

    We found that most neurons are polysemantic meaning they activate on many different human concepts. We used “dictionary learning”, borrowed from classical machine learning, to find monosemantic features from polysemantic neurons.

    https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html

    Comments