We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.

We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.

A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.

  • Artisian@lemmy.worldOP
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    3
    ·
    2 days ago

    Could you flesh this out for me, I’m not sure I understand? I think you’re saying:

    1. experts are not well posed to catch biases and untruths generated by genAI in their (research work)
    2. because as an academic climate scientist, your day-to-day work is spent on new climate science (and not established climate science).

    I don’t see why (2) implies (1), but I agree with (2). As I intended it, (1) is my main claim. I follow it up with

    1. if an expert cannot spot a bias or flaw in genAI output, then they wouldn’t catch it from a peer either.

    I don’t see how (2) helps with (3) either.

    • naught101@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      1 day ago

      GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context). But it also has noise added, so some variance comes back, but there’s no guaranteed that that mean+noise produces something meaningful/true/valuable.

      But, GenAI is very good at producing syntactically correct language. This is a problem because it lulls the reader into a sense that the author knows what it’s talking about, when it doesn’t. When a junior scientist produces text, the awkward language alerts the reviewer to the poor thinking (same with a junior coder producing weak code). With an LLM, you don’t get that - it produces the impression of knowledge without any actual understanding.

      This, combined with the lure of efficiency and the feeling of effectiveness, make it a honey pot for quick but sloppy thinking. I think this is what connects 2 to 1, especially in domains where it’s hard to get external validation from other people who can understand what you’re trying to mean, and not just say.

      As for your point 3: I agree. People don’t spot flaws already, and that’s part of why we saw the replication crisis in behavioural science. Adding LLMs to the mix will just make things like that more likely.

      • Artisian@lemmy.worldOP
        link
        fedilink
        English
        arrow-up
        1
        ·
        15 hours ago

        (There is a separate claim about genAI producing the average/typical behavior from its inputs; this isn’t true for RLHF and boosting reasons, have a gander at boosting and wisdom of the crowds for how you can take a cheap source of mediocrity and create something much better. It is true that genAI outputs feel unoriginal and omnipresent, as so very many folks are throwing around genAI outputs everywhere. But I think that’s not a property of the statistics, rather of the economics/behavioral science.)

        • naught101@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          12 hours ago

          RLHF and boosting

          I don’t really see how those avoid the problem. I can see how they produce much better training results, but ultimately a trained LLM is a neural net that has fixed set of parameters that represent a function over n-dimensional space (I.e. the information in the context window), plus a noise term. If you turn the temperature down to zero, they are perfectly deterministic, no? That’s the mean that they are producing values around…

          • Artisian@lemmy.worldOP
            link
            fedilink
            English
            arrow-up
            1
            ·
            12 hours ago

            GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context).

            Perhaps this sentence is not what you wanted then? The mean of whatever it’s trained on (that is, the training data) need not be the mean of the trained LLM function. This is because we use complicated loss functions to update our LLMs, stochastic processes in training, and that we change those functions after training on input data by using RLHF and other models (eg, via using a mixture of itself to get better outputs with the boosting algorithm).

            You are right that once we fix an LLM’s parameters, at temp 0 they are deterministic, though they are recursive functions so it’s weird to talk about a long string of outputs at temp 0 as the ‘mean’ of the LLM; high temp outputs won’t ‘regress’ to this temp 0 output over time, for example. There also may be hallucinations or mistakes that are common at temp 0, but extremely rare at all other temperatures (because these mistakes only get made when it looks at its own temp 0 context). So I’m not sure the deterministic output string(s) are particularly useful.

            (An analogy that I suspect isn’t very good: the butterfly effect, adding a single small air perturbation can lead to drastically different deterministic weather simulation results. This is very much the norm for LLM outputs)

            (They are somewhat more than a neural network; the ‘attention’ layers add something pretty strange. And output temperature isn’t implemented as a noise term, but this is a fine way to think of it.)

      • Artisian@lemmy.worldOP
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        1
        ·
        edit-2
        15 hours ago

        So I think you’ve landed on the problem that I have with student use, and that the article has with scientist default use. It is very bad for folks to use AI to try to do the knowledge work directly, yet people are making this mistake constantly. YSK: it’s much better when you make AI increase the friction of knowledge work. (note neither are claiming that this use is particularly good; we’re doing damage control.)

        I definitely notice sloppy thinking when I see it in someone else’s review of my papers. You certainly pick it up when you read climate-change-denialists. Experts are great at noticing sloppy thinking when it disagrees with them. That is the relationship you want with genAI if you are using it for knowledge work. Make it disagree with you, so you notice where it is very sloppy (and sometimes where you’ve been sloppy because the bad thinker is kinda right).

        • naught101@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          13 hours ago

          Yeah. Agree in theory. I’m a bit sus on the adversarial AI approach in practice, because most LLMs seem to produce a lot of sycophancy, and I think that’s a far too easy trap to fall in to.

          I do think there are some practical uses, such as reducing availability bias by asking it “is there anything I’ve forgotten to consider in this thing I wrote?” AFTER you’ve done the thinking and writing. But I think the temptation to use it for generative purposes is pretty huge, and likely to suck a lot of people in.