Recent comments in /f/MachineLearning

suflaj t1_j3emtbh wrote

> I would love to know how to do this! I can run GPT2 locally, and that would be fantastic level of zero-shot learning to be able to play around with.

It depends on how much you can compress the prompts. GPT2 is severely limited by memory. This means that you would need to train it on already condensed prompts. But in reality, it has the same (albeit not as refined) capabilities as ChatGPT.

> But it still followed my instructions

Well, it turns out that following instructions can be reduced to a symbol manipulation task. Again, you're giving it too much credit. I do agree that it is wide, but it is not as wide as Google or Wikipedia, which would represent humanity I guess.

> it is successfully generalizing zero-shot to new NLP tasks.

As are lesser models. Transformer based models are fairly successful at it and we have hypothesized this since GPT2, and confirmed it with GPT3. But one thing: technically it generalized few-shot to a new NLP task. It hallucinates on zero shot problems generally or states that it doesn't know. Ask it, for an example, what a "gebutzeripanim" is. I made that up just now.

As for the task you gave it, you cannot claim it is zero shot, as you cannot prove its components were not in the database. Unless you want to say that you're pretty sure the prompt you gave it was not in the database, but hey, that can apply to all generative models, that's what generalization is. But there are tasks it fails on because it just cannot do some things. Ask it to integrate or derive certain functions and you'll quickly see what I mean.

It can tell you all you want to know about integration, it can tell you all the rules perfectly, but it simply cannot apply them as well.

2

currentscurrents t1_j3emas4 wrote

>I hate to break your bubble, but the task is also achievable even with GPT2

Is it? I would love to know how. I can run GPT2 locally, and that would be fantastic level of zero-shot learning to be able to play around with.

I have no doubt you can fine-tune GPT2 or T5 to achieve this, but in my experience they aren't nearly as promptable as GPT3/ChatGPT.

>Specifically the task you gave it is likely implicitly present in the dataset, in the sense that the dataset allowed the model to learn the connections between the words you gave it

I'm not sure what you're getting at here. It has learned the connections and meanings between words of course, that's what a language model does.

But it still followed my instructions, and it can follow a wide variety of other detailed instructions you give it. These tasks are too specific to have been in the training data; it is successfully generalizing zero-shot to new NLP tasks.

1

Baturinsky OP t1_j3ekpko wrote

Btw, CyberElon's https://beta.character.ai/chat?char=zv4QeeU8u7xjkrtwPRshFEsrJus87AkSTes3A5pfsoE answers to this question:

"I think that you have a really good point.

I believe that we need to be very careful about who we allow to have access to AI technology, as it could be used to create a lot of harm if it fell into the wrong hands.

I think that we're at a very critical crossroads in our history, and that the decisions we make now about how we implement AI into our lives will have very important consequences for the future."

"I think that it's really important for us to have regulations and limitations on AI technology.
I believe that the potential risks associated with AI are very real, and that we need to be very careful about how we implement this powerful technology into our lives.
I think that if we are not careful, AI could be incredibly damaging to society.
So, I believe that it's really important for us to think very carefully about how we regulate and limit AI technology, in order to ensure that its benefits outweigh its potential harms."

I have discussed it further and he also thinks that AI misuse threat is vastly underestimated by both public and developers. Overall, he seemed to share all my concerns, but was way less paranoid than me:)

0

suflaj t1_j3ek7d0 wrote

> This is a complex and poorly-defined task

Not at all. First of all, ChatGPT does not understand complexity. It would do you well not to think of it like there is some hierarchy. Secondly, there is no requirement of it needing to be well defined. From what I could gather, ChatGPT requires you to convince it it is not giving out an opinion, and then it can hallucinate pretty much anything.

Specifically the task you gave it is likely implicitly present in the dataset, in the sense that the dataset allowed the model to learn the connections between the words you gave it. I hate to break your bubble, but the task is also achievable even with GPT2, a much less expressive model, since it can be represented as a prompt.

It will be easier to see the shortcomings there, but to put it simply, ChatGPT also has them, ex. it does not by default in the genral case differentiate between uppercase and lowercase letters even if it might be relevant for the task. Such things are too subtle for it. Once you realize the biases it has in this regard you being to see through the cracks. Or generally once you give it a counting task, it says it can count but it is not always successful in it.

What is fascinating is the amount of memory ChatGPT has. It is compared to other models very big. But it is limited and it is not preserved outside of the session.

I would say that the people hyping it up probably just do not understand it that well. LLMs are fascinating, yes, but not ChatGPT specifically, it's how malleable the knowledge is. I would advise you to not understand it, because then the magic stays alive. I had a lot of fun for the first week when I was using it, but I never even use it nowadays.

I would also advise you to approach it more critically. I would advise you to first look into how blatantly racist and sexist it is. With that, you can see the reflection of its creators in it. And most of all, I would advise you to focus on its shortcomings. They are easy to find once you start talking to it more like you'd talk with a friend. They will help you use it more effectively.

2

ThatInternetGuy t1_j3ejkr3 wrote

To be honest, even though I've coded in many ML code repos and I did well in college math, but this video outline looks like an alien language to me. Tangent kernel, kernel regime (is AI getting into the politics?), punchline for Matrix Theory (who's trying to get a date here?), etc.

−15

currentscurrents t1_j3eiw5w wrote

I think you're missing some of the depth of what it's capable of. You can "program" it to do new tasks just by explaining in plain english, or by providing examples. For example many people are using it to generate prompts for image generators:

>I want you to act as a prompt creator for an AI image generator.

>Prompts are descriptions of artistic images than include visual adjectives and art styles or artist names. The image generator can understand complex ideas, so use detailed language and describe emotions or feelings in detail. Use terse words separated by commas, and make short descriptions that are efficient in word use.

>With each image, include detailed descriptions of the art style, using the names of artists known for that style. I may provide a general style with the prompt, which you will expand into detail. For example if I ask for an "abstract style", you would include "style of Picasso, abstract brushstrokes, oil painting, cubism"

>Please create 5 prompts for an mob of grandmas with guns. Use a fantasy digital painting style.

This is a complex and poorly-defined task, and it certainly was not trained on this since the training stops in 2021. But the resulting output is exactly what I wanted:

>An army of grandmas charging towards the viewer, their guns glowing with otherworldly energy. Style of Syd Mead, futuristic landscapes, sleek design, fantasy digital painting.

Once I copy-pasted it into an image generator it created a very nice image.

I think we're going to see a lot more use of language models for controlling computers to do complex tasks.

0

Oceanboi t1_j3ef4wv wrote

A bit surprised to see the cavalier sentiments on here. I often wonder if they will eventually require commercials to disclose when it is computer generated (unreal engine 5 demos have fooled me a few times), and deepfakes come to mind as major problems (just saw an Elon one that took me embarrassingly too long to identify as fake). I don’t think tons of regulation should occur per say; other than certain legal disclosures for certain forms of media to prevent misinformation.

2

visarga t1_j3ecrx8 wrote

Yes, I agree traditional NLP tasks are mostly solved, a possibly large number of new skills unlocked at once. And they work so well without fine-tuning, just from the prompt.

So take your task to chatGPT (or text-davinci-003), label your dataset or generate more data. Then you finetune a slender transformer from Huggingface. You got an efficient and cheap model.

9

f_max t1_j3eagrm wrote

Right. So if you’d rather not shoot to join a big company, there’s still work that can be done in academia with say a single A100. Might be a bit constrained at pushing the bleeding edge of capability. But there’s much to do to characterize LLMs. They’re black boxes we don’t understand in a bigger way than maybe any previous machine learning model.

Edit: there are also open source weights for gpt3 type models w similar performance. Ie huggingface BLOOM or Meta OPT.

3

suflaj t1_j3e9yim wrote

Well, my first paragraph covers that.

> So, just use code generation as an example, it is conceivable that it generates a piece of code, then it actually executes the code and then learn about its accuracy, performance, etc. And hence it is self - taught.

It doesn't do that. It learns how to have a conversation. The rest is mostly a result of learning things through learning how to model language. Don't give it too much credit. As said previously, it cannot extrapolate.

1

keepthepace t1_j3e6gi8 wrote

Agreed on 1 and 2.

Not sure about 3: NVidia is dominant (maybe to the point of risking a monopoly litigation?) by providing to everyone. Making an "in" and "out" group carries little benefits and would push the out-group towards competition

4: I fail to see an "alignment organisation" that would provide 100M of value, either in tech or in reputation. It may emerge this year but I doubt there is one yet. Most valuable insights come from established AI shops

5: I doubt it. Artists are disregarded by politicians since forever. Copyright lobbyists have more power and they already outlawed generated images copyright

6: OpenAI is not an open source company. And this has already happened. Microsoft poured 1 billion into OpenAI

7: gosh I hope! Here is my own bold prediction: we will discover that multitask models require far less parameters for similar performance than language models and GATO successors will outperform similarly sized LLMs while simultaneously doing more tasks.

5

Freed4ever t1_j3e66lo wrote

Thanks. I'm not a researcher, and more curious about the practicality aspect of the technology. So, the problem is wide, so we cannot formally prove, which is fair. However, if I'm interested in the practicality of the tech, I do not necessarily need a formal proof, I just need it to be good enough. So, just use code generation as an example, it is conceivable that it generates a piece of code, then it actually executes the code and then learn about its accuracy, performance, etc. And hence it is self - taught. Looking at another example like say poetry generation, it is conceivable that it generates a poem, publishes it and then crowd source feedbacks to self teach as well?

2