1
00:00:11,119 --> 00:00:12,718
really hard. It took me like 10 its with

2
00:00:12,718 --> 00:00:16,960
Gemini to have Harsha riding a lobster.

3
00:00:16,960 --> 00:00:18,480
Gemini didn't want to do it, but we

4
00:00:18,480 --> 00:00:20,960
figured it out. Okay, welcome to harness

5
00:00:20,960 --> 00:00:24,399
night. Um, first order of business. Uh,

6
00:00:24,399 --> 00:00:26,800
how does everyone like the new YC paper

7
00:00:26,800 --> 00:00:29,199
club look?

8
00:00:29,199 --> 00:00:31,839
Well, this is like an idea that I kind

9
00:00:31,839 --> 00:00:34,719
of blurted out at EV, our head of design

10
00:00:34,719 --> 00:00:37,600
at YC, and two days later, she came back

11
00:00:37,600 --> 00:00:38,640
with this, and I'm like, "This is

12
00:00:38,640 --> 00:00:40,799
amazing." So, Ev in the back there,

13
00:00:40,799 --> 00:00:43,884
please take a bow.

14
00:00:43,884 --> 00:00:45,679
[applause]

15
00:00:45,679 --> 00:00:49,359
Very exciting. All right, why harnesses?

16
00:00:49,359 --> 00:00:51,600
Um, I mean, it's just a rapper. This is

17
00:00:51,600 --> 00:00:53,520
just scaffolding. This is just like

18
00:00:53,520 --> 00:00:56,238
prompt engineering. Um, why would this

19
00:00:56,238 --> 00:00:59,119
be at all worthy of a night? Uh, here's

20
00:00:59,119 --> 00:01:00,878
a great little Reddit that was only a

21
00:01:00,878 --> 00:01:03,119
month ago, which is actually uh the most

22
00:01:03,119 --> 00:01:04,879
aggressive. I'm not sure this kind of

23
00:01:04,879 --> 00:01:06,799
prompt engineering belongs at a top tier

24
00:01:06,799 --> 00:01:09,438
machine learning conference. Um, here's

25
00:01:09,438 --> 00:01:11,438
another great one. Uh, context

26
00:01:11,438 --> 00:01:14,159
engineering is not a research problem.

27
00:01:14,159 --> 00:01:17,519
Um, so I think harnesses have long been

28
00:01:17,519 --> 00:01:22,799
belittled uh as um subpar research. yet

29
00:01:22,799 --> 00:01:26,239
it is uh literally gives us an 18% bump

30
00:01:26,239 --> 00:01:28,000
in the difference between harness one

31
00:01:28,000 --> 00:01:31,438
and harness two. Um, and as Seth will

32
00:01:31,438 --> 00:01:32,799
tell us is the difference between

33
00:01:32,799 --> 00:01:35,759
getting ArcGI to work and not. And so

34
00:01:35,759 --> 00:01:38,640
this is a uh obviously worthy of some

35
00:01:38,640 --> 00:01:40,959
some amount of research. And I think if

36
00:01:40,959 --> 00:01:43,118
you look at this the classic meter plots

37
00:01:43,118 --> 00:01:46,319
release date and how long uh an agent

38
00:01:46,319 --> 00:01:49,040
can be running. Um, literally we go a

39
00:01:49,040 --> 00:01:50,640
lot of this progress has been because of

40
00:01:50,640 --> 00:01:52,560
harnesses. And so I call this the static

41
00:01:52,560 --> 00:01:54,799
harness era where there isn't

42
00:01:54,799 --> 00:01:56,399
self-improvement on the harness. And

43
00:01:56,399 --> 00:01:59,359
then this later latest latest um maybe

44
00:01:59,359 --> 00:02:00,959
the last 6 months has all been on the

45
00:02:00,959 --> 00:02:02,799
self improving harnesses and think and

46
00:02:02,799 --> 00:02:05,438
we'll get through all these and it's I

47
00:02:05,438 --> 00:02:07,759
actually saw this plot in a presentation

48
00:02:07,759 --> 00:02:09,919
by the CEO of trajectory. I actually

49
00:02:09,919 --> 00:02:12,318
really liked it. was talking about how

50
00:02:12,318 --> 00:02:16,159
the we we keep measuring perplexity and

51
00:02:16,159 --> 00:02:17,919
you know this somewhat correlates to IQ

52
00:02:17,919 --> 00:02:19,840
and and how intelligent the model is and

53
00:02:19,840 --> 00:02:21,840
so we keep pushing more and more up this

54
00:02:21,840 --> 00:02:25,759
intelligence uh IQ dimension but we're

55
00:02:25,759 --> 00:02:28,800
not leveraging test time experience very

56
00:02:28,800 --> 00:02:30,479
much so we're generating all this test

57
00:02:30,479 --> 00:02:32,719
time experience we have a new domain uh

58
00:02:32,719 --> 00:02:34,878
and it does and we don't really quickly

59
00:02:34,878 --> 00:02:36,719
adapt and if you remember one of the

60
00:02:36,719 --> 00:02:39,039
first YC paper clubs we did um I've been

61
00:02:39,039 --> 00:02:42,159
doing this experiment where um the as I

62
00:02:42,159 --> 00:02:44,479
increase the number of samples online

63
00:02:44,479 --> 00:02:45,919
how do you actually learn from just

64
00:02:45,919 --> 00:02:47,519
batch size one we don't really don't

65
00:02:47,519 --> 00:02:50,318
have that structure we have ICL and then

66
00:02:50,318 --> 00:02:52,878
once ICL gets saturated after just

67
00:02:52,878 --> 00:02:54,959
loaded like 40 or 50 it doesn't actually

68
00:02:54,959 --> 00:02:57,519
improve um on the valet at all then you

69
00:02:57,519 --> 00:02:59,519
have to go to Laura small rank then you

70
00:02:59,519 --> 00:03:02,000
go to Laura big rank then you go to SFT

71
00:03:02,000 --> 00:03:03,598
and it's kind of weird that we have this

72
00:03:03,598 --> 00:03:05,919
like different training procedures and

73
00:03:05,919 --> 00:03:08,000
so I think that's really where harnesses

74
00:03:08,000 --> 00:03:10,318
are shining right

75
00:03:10,318 --> 00:03:12,639
And what Arc AGI exposes, that's kind of

76
00:03:12,639 --> 00:03:16,719
the main uh um uh point of ARGI is how

77
00:03:16,719 --> 00:03:20,719
quickly it adapts to uh a new problem, a

78
00:03:20,719 --> 00:03:23,360
new a new distribution and does well in

79
00:03:23,360 --> 00:03:26,719
it. And so, um, ArcGI actually went

80
00:03:26,719 --> 00:03:28,959
through the batch with me, winter 26,

81
00:03:28,959 --> 00:03:30,639
and we helped Greg, you know, look at

82
00:03:30,639 --> 00:03:32,000
this and like the amount of thought and

83
00:03:32,000 --> 00:03:33,439
attention, as I mentioned last time,

84
00:03:33,439 --> 00:03:35,759
that goes into these these these games

85
00:03:35,759 --> 00:03:37,360
to make sure they're all orthogonal

86
00:03:37,360 --> 00:03:39,439
skills from game one to game two, so

87
00:03:39,439 --> 00:03:40,878
that it's isolating this fluid

88
00:03:40,878 --> 00:03:44,400
intelligence measure. Um really uh

89
00:03:44,400 --> 00:03:47,120
Claude uh Opus was one of the the first

90
00:03:47,120 --> 00:03:49,919
that was actually verified um on the the

91
00:03:49,919 --> 00:03:51,598
hold out on the private that no one else

92
00:03:51,598 --> 00:03:53,439
has access to other than Greg and

93
00:03:53,439 --> 00:03:55,519
Chalet. Um and the best that they got

94
00:03:55,519 --> 00:03:59,519
was 30%. And just with um some harness

95
00:03:59,519 --> 00:04:01,840
uh this thing that doesn't deserve any

96
00:04:01,840 --> 00:04:03,438
research just some wrapper and some

97
00:04:03,438 --> 00:04:06,400
scaffolding we can get to 95 and AVO

98
00:04:06,400 --> 00:04:08,719
from Nvidia got to 100%. So, Prime Agent

99
00:04:08,719 --> 00:04:11,199
and Nvidia, which both recently just

100
00:04:11,199 --> 00:04:12,878
came out.

101
00:04:12,878 --> 00:04:14,719
And the other thing I want to add, so

102
00:04:14,719 --> 00:04:17,680
when uh Carpathy launched his auto

103
00:04:17,680 --> 00:04:20,238
researcher thing in March, I want to say

104
00:04:20,238 --> 00:04:23,199
it was um I forked it and I was playing

105
00:04:23,199 --> 00:04:24,319
around with it. And all I wanted to do

106
00:04:24,319 --> 00:04:25,839
was make like a little user interface to

107
00:04:25,839 --> 00:04:27,120
kind of see what's happening and track

108
00:04:27,120 --> 00:04:30,639
it and and I ended up building a harness

109
00:04:30,639 --> 00:04:32,399
by accident. I didn't mean to, but it

110
00:04:32,399 --> 00:04:34,560
was just like I wanted to see it. And

111
00:04:34,560 --> 00:04:37,439
and basically what it is is you specify

112
00:04:37,439 --> 00:04:39,918
a purpose. And in this example, which is

113
00:04:39,918 --> 00:04:41,439
actually a true one I gave, is like

114
00:04:41,439 --> 00:04:45,600
diffusion LM don't beat ARLM. But maybe

115
00:04:45,600 --> 00:04:48,160
if I ensemble, if I shard the diffusion

116
00:04:48,160 --> 00:04:49,839
LM into a bunch of different ones

117
00:04:49,839 --> 00:04:51,439
because there's such high arithmetic

118
00:04:51,439 --> 00:04:53,600
intensity per GPU on a diffusion model

119
00:04:53,600 --> 00:04:57,120
versus AR that I can actually um in

120
00:04:57,120 --> 00:04:58,800
aggregate by sharding them, I can get

121
00:04:58,800 --> 00:05:00,478
better better results. And so I just

122
00:05:00,478 --> 00:05:02,560
give this to a per as the purpose to the

123
00:05:02,560 --> 00:05:07,759
to my uh uh to my suite of agents um my

124
00:05:07,759 --> 00:05:09,279
swarm of agents. I give it some seed

125
00:05:09,279 --> 00:05:12,240
ideas. I want to vary the ensemble size

126
00:05:12,240 --> 00:05:14,959
shard at different amounts

127
00:05:14,959 --> 00:05:16,478
100 times, 10 times, five times,

128
00:05:16,478 --> 00:05:19,120
whatever. Um and then I specify a

129
00:05:19,120 --> 00:05:21,680
valometric. Um and maybe I want to do

130
00:05:21,680 --> 00:05:24,560
GSMAK and like GPT2 setup or something

131
00:05:24,560 --> 00:05:26,879
like that. And then I have a scoping

132
00:05:26,879 --> 00:05:29,680
agent that will kind of look up papers

133
00:05:29,680 --> 00:05:31,439
that have that are are similar, look up

134
00:05:31,439 --> 00:05:34,399
GitHub repos. Um, we'll give it to a PI

135
00:05:34,399 --> 00:05:37,918
agent named Chris Ray. Um, who will then

136
00:05:37,918 --> 00:05:40,560
uh give it to a research agent um named

137
00:05:40,560 --> 00:05:44,079
John South Khan. Uh, there he is. Um,

138
00:05:44,079 --> 00:05:45,918
and then he'll work on it and then he'll

139
00:05:45,918 --> 00:05:47,600
start doing some stuff, give it to a

140
00:05:47,600 --> 00:05:49,439
council for some help. That's where I

141
00:05:49,439 --> 00:05:51,279
come in and me and Yaso will give you

142
00:05:51,279 --> 00:05:52,879
some feedback and then you work on it a

143
00:05:52,879 --> 00:05:54,800
bit more and then whenever you're ready

144
00:05:54,800 --> 00:05:56,879
and Chris Ray will kind of keep tabs on

145
00:05:56,879 --> 00:05:58,639
you, keep nagging you every every hour

146
00:05:58,639 --> 00:06:00,959
or so then it goes to this author agent

147
00:06:00,959 --> 00:06:03,279
to say okay now start stop freeze the

148
00:06:03,279 --> 00:06:06,720
idea start writing ablations and um and

149
00:06:06,720 --> 00:06:08,399
then start writing the paper and so then

150
00:06:08,399 --> 00:06:10,240
what this has turned into so this is the

151
00:06:10,240 --> 00:06:12,639
scoping agent uh and then you have this

152
00:06:12,639 --> 00:06:14,800
cockpit that kind of you can view from

153
00:06:14,800 --> 00:06:17,680
anywhere we do the uh tail scale up so

154
00:06:17,680 --> 00:06:19,199
you can actually view this URL from

155
00:06:19,199 --> 00:06:21,120
anywhere and see how it's progressing

156
00:06:21,120 --> 00:06:24,079
and and talk to it right there. Um, it

157
00:06:24,079 --> 00:06:26,478
just sends you email updates if you

158
00:06:26,478 --> 00:06:28,478
build the whole 1B thing in here. And

159
00:06:28,478 --> 00:06:29,839
then it actually starts publishing some

160
00:06:29,839 --> 00:06:31,600
papers and like I've read the papers.

161
00:06:31,600 --> 00:06:33,038
They're actually like good. They started

162
00:06:33,038 --> 00:06:35,360
out in March, April like not so good

163
00:06:35,360 --> 00:06:37,519
like you know May or whatever. But and

164
00:06:37,519 --> 00:06:40,399
now I just basically give eight ideas to

165
00:06:40,399 --> 00:06:42,879
eight H100 nodes and each one has eight

166
00:06:42,879 --> 00:06:44,879
eight H100s. and just keep going. And

167
00:06:44,879 --> 00:06:46,720
then I check in and they give give me

168
00:06:46,720 --> 00:06:48,639
back papers. And it's kind of wild what

169
00:06:48,639 --> 00:06:50,079
we're what we're dealing with there. And

170
00:06:50,079 --> 00:06:51,759
what's changed really is just the

171
00:06:51,759 --> 00:06:53,279
scaffolding that you put on top of it to

172
00:06:53,279 --> 00:06:54,879
allow it. And I I didn't spend an

173
00:06:54,879 --> 00:06:56,319
enormous amount of time on this, but

174
00:06:56,319 --> 00:06:58,000
this is largely how I I do a lot of my

175
00:06:58,000 --> 00:07:00,319
research, at least the the initial idea.

176
00:07:00,319 --> 00:07:03,120
Um, and so yeah, so I'll just give this

177
00:07:03,120 --> 00:07:04,720
six different ideas of things I want to

178
00:07:04,720 --> 00:07:06,720
try and just leave it alone and let it

179
00:07:06,720 --> 00:07:08,560
rip. And so the things that we can now

180
00:07:08,560 --> 00:07:10,240
do just because of harnesses on the same

181
00:07:10,240 --> 00:07:14,720
exact weight file is just wild. Um, so I

182
00:07:14,720 --> 00:07:16,560
looked on I spent the the weekend just

183
00:07:16,560 --> 00:07:19,439
in prep for this reading um a whole

184
00:07:19,439 --> 00:07:22,478
bunch of the classic literature from

185
00:07:22,478 --> 00:07:24,800
self-refined to reflection to Voyager to

186
00:07:24,800 --> 00:07:27,439
all the tool former and I just wanted to

187
00:07:27,439 --> 00:07:29,759
just do like a my my shot at a five

188
00:07:29,759 --> 00:07:32,560
minute like history how did we get here?

189
00:07:32,560 --> 00:07:34,639
Um I think it was it's kind of important

190
00:07:34,639 --> 00:07:36,160
and I don't think a lot of people that

191
00:07:36,160 --> 00:07:38,478
are entering AI now kind of like have

192
00:07:38,478 --> 00:07:40,079
the context. So this is not going to be

193
00:07:40,079 --> 00:07:41,360
in chronological order. It actually

194
00:07:41,360 --> 00:07:43,439
doesn't make sense to to teach it that

195
00:07:43,439 --> 00:07:45,598
way or show it that way. I'm not I'm

196
00:07:45,598 --> 00:07:47,598
going to skip a lot of papers and I may

197
00:07:47,598 --> 00:07:49,360
course grain the papers absurdly. So,

198
00:07:49,360 --> 00:07:51,439
apologies. So, the initial harness the

199
00:07:51,439 --> 00:07:55,839
GPT2 February 2019 is a while not end of

200
00:07:55,839 --> 00:07:58,478
end of sequence loop. That's and then we

201
00:07:58,478 --> 00:08:00,720
basically have top P sampling and then

202
00:08:00,720 --> 00:08:02,560
we have the environment and that's the

203
00:08:02,560 --> 00:08:03,759
harness. And there's really not that

204
00:08:03,759 --> 00:08:05,680
much. There's no tool calling, there's

205
00:08:05,680 --> 00:08:07,360
no skills, there's nothing like that.

206
00:08:07,360 --> 00:08:10,478
And so that's the V0ero harness.

207
00:08:10,478 --> 00:08:12,720
And so basically if you have this GSMAK

208
00:08:12,720 --> 00:08:14,160
example, you're a math, remember the

209
00:08:14,160 --> 00:08:15,918
system prompt. You are a math teacher.

210
00:08:15,918 --> 00:08:17,519
This is like the persona stuff we used

211
00:08:17,519 --> 00:08:19,680
to have. Um there's a context. Susie has

212
00:08:19,680 --> 00:08:21,120
five bucks, she spends three. How much

213
00:08:21,120 --> 00:08:23,199
does she have now? And then there's just

214
00:08:23,199 --> 00:08:24,720
no chain of thought. It was just like

215
00:08:24,720 --> 00:08:27,120
four hashes and two end of state and end

216
00:08:27,120 --> 00:08:30,639
of sequence that is uh the GS MK format.

217
00:08:30,639 --> 00:08:33,038
We measure uh what happens after the

218
00:08:33,038 --> 00:08:35,200
four hashes. We get accuracy and you get

219
00:08:35,200 --> 00:08:37,918
a plus one or minus one if it's wrong.

220
00:08:37,918 --> 00:08:40,799
And then the entire last six years has

221
00:08:40,799 --> 00:08:42,479
been giving more functionality into the

222
00:08:42,479 --> 00:08:45,278
harness into a static harness. And so we

223
00:08:45,278 --> 00:08:47,360
said okay well what if we give that back

224
00:08:47,360 --> 00:08:49,120
into the context a bunch of examples

225
00:08:49,120 --> 00:08:51,278
like that like well here's an example

226
00:08:51,278 --> 00:08:52,799
and let's say now we'll switch it to

227
00:08:52,799 --> 00:08:56,879
four and one and hopefully we'll have um

228
00:08:56,879 --> 00:08:58,879
from that previous example we'll say oh

229
00:08:58,879 --> 00:09:01,120
okay this makes sense I can will help to

230
00:09:01,120 --> 00:09:02,720
learn and actually this is with a fshot

231
00:09:02,720 --> 00:09:05,759
learners paper um back in July in 2020

232
00:09:05,759 --> 00:09:07,278
and then we had chain of thought and

233
00:09:07,278 --> 00:09:10,080
says well directly predicting hash hash

234
00:09:10,080 --> 00:09:12,399
2 is maybe difficult what we'll do is

235
00:09:12,399 --> 00:09:15,200
we'll smear the the compute the logic

236
00:09:15,200 --> 00:09:17,679
over many more tokens and we'll train it

237
00:09:17,679 --> 00:09:19,919
to get to that two now rather than just

238
00:09:19,919 --> 00:09:22,080
output the two and that was the chain of

239
00:09:22,080 --> 00:09:25,039
thought idea uh very cool and that was

240
00:09:25,039 --> 00:09:27,839
all context innovations uh and output

241
00:09:27,839 --> 00:09:29,120
space innovations action space

242
00:09:29,120 --> 00:09:31,278
innovations um and then then we came up

243
00:09:31,278 --> 00:09:33,919
with these tool former and webgbt webgbt

244
00:09:33,919 --> 00:09:35,759
actually came out first then tool former

245
00:09:35,759 --> 00:09:37,679
which is the idea of giving it tools and

246
00:09:37,679 --> 00:09:39,600
it can call a tool that's just really

247
00:09:39,600 --> 00:09:41,679
just a JSON object with a bunch of

248
00:09:41,679 --> 00:09:43,519
things specified but in this example

249
00:09:43,519 --> 00:09:45,759
just let's say subtraction instead of me

250
00:09:45,759 --> 00:09:48,320
calculating in the weight file what is 5

251
00:09:48,320 --> 00:09:51,600
- 2 I can just call Python and call sub

252
00:09:51,600 --> 00:09:54,320
53 and it'll tell me two and so that's

253
00:09:54,320 --> 00:09:56,639
pretty cool and you can expose all the

254
00:09:56,639 --> 00:09:59,200
tools in the system prompt and that's

255
00:09:59,200 --> 00:10:01,120
where tools come from and then we

256
00:10:01,120 --> 00:10:03,120
figured out megpt which is big the one

257
00:10:03,120 --> 00:10:04,720
of the coolest tools is being able to

258
00:10:04,720 --> 00:10:07,200
read and write to your own context and

259
00:10:07,200 --> 00:10:08,799
so before then all we could do is just

260
00:10:08,799 --> 00:10:10,559
append append append append to the

261
00:10:10,559 --> 00:10:12,159
context now we said what if I actually

262
00:10:12,159 --> 00:10:14,958
give you create read, update, delete on

263
00:10:14,958 --> 00:10:17,200
the context itself and we'll se separate

264
00:10:17,200 --> 00:10:18,639
just this one little chunk called a

265
00:10:18,639 --> 00:10:22,559
memory that you'll be able to to update.

266
00:10:22,559 --> 00:10:24,720
And then Voyager said, okay, well, we

267
00:10:24,720 --> 00:10:26,480
have these tools, but like what if I

268
00:10:26,480 --> 00:10:27,919
want to chain together these tools to

269
00:10:27,919 --> 00:10:30,240
achieve a task and then I learn it and

270
00:10:30,240 --> 00:10:32,399
how do I distill it back into the the

271
00:10:32,399 --> 00:10:35,120
system prompt to make it learn forever?

272
00:10:35,120 --> 00:10:37,278
And this is the skill this where skills

273
00:10:37,278 --> 00:10:40,078
kind of came about. And uh they did this

274
00:10:40,078 --> 00:10:42,159
on Minecraft. Um and this the Voyager

275
00:10:42,159 --> 00:10:44,078
paper and it's very cool paper. This is

276
00:10:44,078 --> 00:10:46,159
like largely now what a skill is. And so

277
00:10:46,159 --> 00:10:49,120
I have this skills.md and I have the

278
00:10:49,120 --> 00:10:51,278
name of it that I can go and search and

279
00:10:51,278 --> 00:10:53,440
here's the procedure. And then uh

280
00:10:53,440 --> 00:10:55,679
intercode this idea of like again in

281
00:10:55,679 --> 00:10:57,440
action space innovation if I can

282
00:10:57,440 --> 00:10:59,759
actually output code uh then I can

283
00:10:59,759 --> 00:11:02,799
basically now I have on the-ly uh tools

284
00:11:02,799 --> 00:11:04,559
or on the fly skills. I actually don't

285
00:11:04,559 --> 00:11:06,799
know if it's a considered a tool because

286
00:11:06,799 --> 00:11:08,879
a tool is technically an API. Am I

287
00:11:08,879 --> 00:11:10,958
outputting a function as a tool or a

288
00:11:10,958 --> 00:11:14,559
skill? I actually am still unsure. Um

289
00:11:14,559 --> 00:11:17,519
and then the react came first then

290
00:11:17,519 --> 00:11:19,839
self-refine then reflection. But this

291
00:11:19,839 --> 00:11:23,039
idea if I have multiple um agents that

292
00:11:23,039 --> 00:11:24,799
that have different roles and they can

293
00:11:24,799 --> 00:11:27,278
help self-improve self-improve on the

294
00:11:27,278 --> 00:11:30,720
context uh then I can um get smarter and

295
00:11:30,720 --> 00:11:32,559
smarter. So, I'll have take some action.

296
00:11:32,559 --> 00:11:35,440
In this example, I I I flipped the three

297
00:11:35,440 --> 00:11:38,720
and the five. Whoops. Um, I sent it to

298
00:11:38,720 --> 00:11:40,720
an internal evaluator to say, is this

299
00:11:40,720 --> 00:11:42,879
right? No, it doesn't look right. I can

300
00:11:42,879 --> 00:11:44,720
either, you know, keep looping here or I

301
00:11:44,720 --> 00:11:46,480
can go to the actual environment to get

302
00:11:46,480 --> 00:11:49,919
back a reward signal. Um, I go here uh

303
00:11:49,919 --> 00:11:51,679
uh to the reflection. They'll say, "Hey,

304
00:11:51,679 --> 00:11:53,679
you actually flip these two, go back,

305
00:11:53,679 --> 00:11:56,720
and I can improve the result." Um and

306
00:11:56,720 --> 00:12:00,078
this is the one of the first like uh uh

307
00:12:00,078 --> 00:12:02,879
ideas of like being multi- aent uh being

308
00:12:02,879 --> 00:12:05,039
let the letting the agent reflect on its

309
00:12:05,039 --> 00:12:08,399
own um output and improve it. And then

310
00:12:08,399 --> 00:12:11,039
this idea of multi- aent goes even

311
00:12:11,039 --> 00:12:13,679
further where I can actually spawn one

312
00:12:13,679 --> 00:12:16,639
of the tools could be I can spawn a a

313
00:12:16,639 --> 00:12:19,519
set of sub aents um and they they

314
00:12:19,519 --> 00:12:21,440
persist in these persistent ripples and

315
00:12:21,440 --> 00:12:22,879
they'll be able to be running and I can

316
00:12:22,879 --> 00:12:25,039
interact with them and I'll have the sub

317
00:12:25,039 --> 00:12:26,958
aent list. I can invoke those ones and I

318
00:12:26,958 --> 00:12:28,958
can keep adding to that the the launch

319
00:12:28,958 --> 00:12:32,240
sub aents and then RLM went even crazier

320
00:12:32,240 --> 00:12:34,240
to allow this in a recursive fashion so

321
00:12:34,240 --> 00:12:36,958
that um and I exposes this RLM query so

322
00:12:36,958 --> 00:12:39,200
I keep uh recursively calling the RLM

323
00:12:39,200 --> 00:12:41,600
query to uh solve a larger class of

324
00:12:41,600 --> 00:12:43,519
problems and that the leafs whenever I

325
00:12:43,519 --> 00:12:45,919
want I can call the LM query to spawn

326
00:12:45,919 --> 00:12:49,120
that LM agent as well and then I have

327
00:12:49,120 --> 00:12:51,039
this main orchestrator agent that's

328
00:12:51,039 --> 00:12:54,159
running all that and that's what I call

329
00:12:54,159 --> 00:12:56,159
like harness v1 one, this whole thing is

330
00:12:56,159 --> 00:12:57,759
like static harnesses. I'm not improving

331
00:12:57,759 --> 00:13:00,240
on the system prompt. I'm not um

332
00:13:00,240 --> 00:13:03,278
updating the harness itself. And this is

333
00:13:03,278 --> 00:13:04,399
this you, of course, we're going to have

334
00:13:04,399 --> 00:13:06,240
to have GStack on the top. He's the one

335
00:13:06,240 --> 00:13:07,360
that lets us do all this. So, we thank

336
00:13:07,360 --> 00:13:10,000
you, Gary. Love GStack. Um but there's

337
00:13:10,000 --> 00:13:11,600
some other ones that deserve a call out

338
00:13:11,600 --> 00:13:14,159
as well. And basically what that means

339
00:13:14,159 --> 00:13:15,759
to summarize all this, you have some

340
00:13:15,759 --> 00:13:18,078
agent spec, you have some system prompt

341
00:13:18,078 --> 00:13:19,919
here. You say, "How many turns am I

342
00:13:19,919 --> 00:13:21,039
allowed? How many tool calls am I

343
00:13:21,039 --> 00:13:22,078
allowed?" You don't want to allow

344
00:13:22,078 --> 00:13:24,399
infinite. You specify a tool list. you

345
00:13:24,399 --> 00:13:26,480
specify a skills list, the sub agent

346
00:13:26,480 --> 00:13:29,679
list, and that's largely what the V1 um

347
00:13:29,679 --> 00:13:32,559
is, and then you put it in a loop. And

348
00:13:32,559 --> 00:13:34,799
this can be spawned on a prompt if I'm

349
00:13:34,799 --> 00:13:37,200
asking it to do something in my Slack

350
00:13:37,200 --> 00:13:39,039
channel, which we'll hear about QM,

351
00:13:39,039 --> 00:13:40,958
which is like I use it every day. It's

352
00:13:40,958 --> 00:13:42,159
the team that made it is here. It's

353
00:13:42,159 --> 00:13:45,039
super exciting. Um and or it's on a

354
00:13:45,039 --> 00:13:46,879
crunch and just wakes up every hour and

355
00:13:46,879 --> 00:13:48,559
it decides to do work just like you

356
00:13:48,559 --> 00:13:50,958
know, anyone else. Um, and so there's

357
00:13:50,958 --> 00:13:52,720
some session management, there's a loop,

358
00:13:52,720 --> 00:13:54,559
there's this context compilation. We're

359
00:13:54,559 --> 00:13:56,240
actually creating the context. I'm

360
00:13:56,240 --> 00:13:59,198
putting uh uh all that into an LM call.

361
00:13:59,198 --> 00:14:02,399
I get back uh the action. And then I I

362
00:14:02,399 --> 00:14:04,240
may or may not have some tools that I

363
00:14:04,240 --> 00:14:07,278
need to invoke and append back into the

364
00:14:07,278 --> 00:14:09,919
context. And that's basically harness

365
00:14:09,919 --> 00:14:12,879
v1. And then the cool part, this is the

366
00:14:12,879 --> 00:14:14,320
most exciting part where we're seeing a

367
00:14:14,320 --> 00:14:16,078
lot of advancements, um, where you're

368
00:14:16,078 --> 00:14:18,639
letting the harness itself learn. either

369
00:14:18,639 --> 00:14:20,159
we're learning the system prompt or

370
00:14:20,159 --> 00:14:22,078
we're learning the harness itself, which

371
00:14:22,078 --> 00:14:24,879
is very trippy. And so, one of the

372
00:14:24,879 --> 00:14:26,720
famous ones that I actually wanted them

373
00:14:26,720 --> 00:14:29,039
to talk, but they're actually running a

374
00:14:29,039 --> 00:14:32,240
150 person DSPY meetup tonight and in in

375
00:14:32,240 --> 00:14:33,600
this in San Francisco. They couldn't

376
00:14:33,600 --> 00:14:35,600
make it. Um, but I've done some podcast

377
00:14:35,600 --> 00:14:37,039
with them before and a super great

378
00:14:37,039 --> 00:14:39,039
community. Is this DSPY? So,

379
00:14:39,039 --> 00:14:42,000
demonstrate, search, uh, predict. uh

380
00:14:42,000 --> 00:14:45,278
they have this idea of of basically uh

381
00:14:45,278 --> 00:14:48,320
taking a train set a small set of

382
00:14:48,320 --> 00:14:50,639
examples and then learning the optimal

383
00:14:50,639 --> 00:14:52,320
system prompt. So I keep iterating,

384
00:14:52,320 --> 00:14:54,320
iterating, iterating. I can't back prop

385
00:14:54,320 --> 00:14:56,320
through that process, but I can do uh

386
00:14:56,320 --> 00:14:57,519
something called genetic programming

387
00:14:57,519 --> 00:14:58,958
where I'm finding candidates, I'm

388
00:14:58,958 --> 00:15:00,720
merging candidates, and with some merge

389
00:15:00,720 --> 00:15:02,480
rule, I'm evaluating, seeing what

390
00:15:02,480 --> 00:15:04,559
happens, and I keep working and working.

391
00:15:04,559 --> 00:15:06,639
And I basically gives me CRUD over the

392
00:15:06,639 --> 00:15:08,639
system prompt itself, and allows me to

393
00:15:08,639 --> 00:15:10,399
choose any system prompt. And then

394
00:15:10,399 --> 00:15:12,639
Darwin machines actually go a step

395
00:15:12,639 --> 00:15:14,320
further. Not only are you allowed to

396
00:15:14,320 --> 00:15:15,600
change the system prompt, but you're

397
00:15:15,600 --> 00:15:17,120
allowed to change the harness itself,

398
00:15:17,120 --> 00:15:19,039
the harness code that is actually

399
00:15:19,039 --> 00:15:21,039
running. And so you can imagine

400
00:15:21,039 --> 00:15:23,039
basically what happens is you have this

401
00:15:23,039 --> 00:15:27,198
archive of um many different uh agents

402
00:15:27,198 --> 00:15:30,159
which is harness and system prompt. Uh

403
00:15:30,159 --> 00:15:33,919
and you sample from them. You push them

404
00:15:33,919 --> 00:15:35,919
uh through you'll actually evaluate how

405
00:15:35,919 --> 00:15:38,240
how they did on some fit fitness

406
00:15:38,240 --> 00:15:41,679
function. Uh you'll uh add it back into

407
00:15:41,679 --> 00:15:43,919
into this archive state. And I skipped

408
00:15:43,919 --> 00:15:45,600
over the the self modify. You can

409
00:15:45,600 --> 00:15:48,159
actually you have a a meta harness that

410
00:15:48,159 --> 00:15:49,839
actually allows the the agent to modify

411
00:15:49,839 --> 00:15:52,159
its own uh harness so that it can become

412
00:15:52,159 --> 00:15:53,839
a different harness and then that

413
00:15:53,839 --> 00:15:56,320
basically loops around loops around and

414
00:15:56,320 --> 00:15:57,919
eventually you get better and better

415
00:15:57,919 --> 00:15:59,600
agents over time. And then the meta

416
00:15:59,600 --> 00:16:01,278
harness which is the main harness is to

417
00:16:01,278 --> 00:16:03,039
produce harnesses, right? Which is a

418
00:16:03,039 --> 00:16:05,519
really meta concept. Um and so this is

419
00:16:05,519 --> 00:16:06,799
the output space where you're

420
00:16:06,799 --> 00:16:08,879
configuring multi- aent context

421
00:16:08,879 --> 00:16:11,440
compilation. Um, and you're you're doing

422
00:16:11,440 --> 00:16:13,360
this and you have allow CRUD on all of

423
00:16:13,360 --> 00:16:15,440
it. And so it just keeps adding more and

424
00:16:15,440 --> 00:16:18,879
more uh CRUD into the harness code, all

425
00:16:18,879 --> 00:16:21,600
the green that you see here, the uh meta

426
00:16:21,600 --> 00:16:23,759
prompt, the system prompts of all the

427
00:16:23,759 --> 00:16:25,759
agents, how many agents there are. Um,

428
00:16:25,759 --> 00:16:28,000
and you kind of grow this uh this meta

429
00:16:28,000 --> 00:16:30,399
harness over time. And then one of the

430
00:16:30,399 --> 00:16:32,958
authors of this paper is the lead author

431
00:16:32,958 --> 00:16:34,320
is actually here tonight. Super

432
00:16:34,320 --> 00:16:35,600
exciting. You got to drive down and talk

433
00:16:35,600 --> 00:16:37,839
to him a lot of it about um continual

434
00:16:37,839 --> 00:16:40,159
harness, which is I love it. And they go

435
00:16:40,159 --> 00:16:43,198
even a step further where one they add

436
00:16:43,198 --> 00:16:46,720
they they add some extra color on the

437
00:16:46,720 --> 00:16:49,039
classes of memory and so they add this

438
00:16:49,039 --> 00:16:50,720
history thing. You know he'll he'll go

439
00:16:50,720 --> 00:16:52,078
into a bunch bunch more detail on that

440
00:16:52,078 --> 00:16:54,639
memory breakdown. But the coolest part I

441
00:16:54,639 --> 00:16:56,799
think is for the classic RL people. I

442
00:16:56,799 --> 00:16:58,480
see Robert back there. He definitely

443
00:16:58,480 --> 00:17:00,799
would would enjoy this Dagger style

444
00:17:00,799 --> 00:17:03,039
online learning where you can actually

445
00:17:03,039 --> 00:17:05,199
update the weight file itself. So you're

446
00:17:05,199 --> 00:17:08,640
actually doing a test time training on

447
00:17:08,640 --> 00:17:10,480
the LLM based on a small amount of

448
00:17:10,480 --> 00:17:12,078
examples that you just learned which I

449
00:17:12,078 --> 00:17:14,240
think is actually a huge huge important

450
00:17:14,240 --> 00:17:16,000
research direction that we should get

451
00:17:16,000 --> 00:17:20,160
working. Anyway, that's it. How you do

452
00:17:20,160 --> 00:17:26,640
all

453
00:17:26,640 --> 00:17:28,880
authors. Uh Ben got food poisoning this

454
00:17:28,880 --> 00:17:30,240
morning so he couldn't make it. It was

455
00:17:30,240 --> 00:17:32,319
very sad but we have three tremendous

456
00:17:32,319 --> 00:17:37,519
authors. uh Seth uh student under shein

457
00:17:37,519 --> 00:17:39,679
is I say it right uh at Princeton a

458
00:17:39,679 --> 00:17:41,279
researcher at Prime Intellect and the

459
00:17:41,279 --> 00:17:44,880
author of Prime Agent um John Sadvalone

460
00:17:44,880 --> 00:17:47,038
who's a the first time we've had a a

461
00:17:47,038 --> 00:17:49,839
call back um pres presenter very excited

462
00:17:49,839 --> 00:17:53,679
about that PhD under Chris an Aelia um

463
00:17:53,679 --> 00:17:56,160
and he's going to be he's the author of

464
00:17:56,160 --> 00:17:59,119
open Jarvis with Ivonica uh from my lab

465
00:17:59,119 --> 00:18:01,759
uh hazy which is a personal uh open

466
00:18:01,759 --> 00:18:02,480
Jarvis

467
00:18:02,480 --> 00:18:03,919
and which I think is really cool. And

468
00:18:03,919 --> 00:18:07,440
then for Josh and Rean who Josh just got

469
00:18:07,440 --> 00:18:09,519
promoted to be head of YC labs which is

470
00:18:09,519 --> 00:18:11,440
really exciting which I think deserves a

471
00:18:11,440 --> 00:18:13,207
little round of applause as well.

472
00:18:13,207 --> 00:18:14,558
[applause]

473
00:18:14,558 --> 00:18:18,000
Um be talking about QM and so QM there

474
00:18:18,000 --> 00:18:20,160
at YC we had this general agent that we

475
00:18:20,160 --> 00:18:21,759
use and they came out with QM a month

476
00:18:21,759 --> 00:18:24,319
ago and it is like meaningfully uh

477
00:18:24,319 --> 00:18:26,160
better and is so functional and I use it

478
00:18:26,160 --> 00:18:28,558
every day. All right, that's all I got.

479
00:18:28,558 --> 00:18:36,558
Thank you so much. [applause]

480
00:18:37,279 --> 00:18:38,960
tonight. Uh, my name is Seth and I'm

481
00:18:38,960 --> 00:18:40,240
going to be presenting Prime Agent,

482
00:18:40,240 --> 00:18:42,880
which is a self-improving ROM harness.

483
00:18:42,880 --> 00:18:45,440
Um, my job here tonight, I think, is to

484
00:18:45,440 --> 00:18:48,319
try to convince all of you to take a

485
00:18:48,319 --> 00:18:52,000
very first principles style approach to

486
00:18:52,000 --> 00:18:53,119
thinking about how you build your

487
00:18:53,119 --> 00:18:54,880
harness. So, we're going to be very,

488
00:18:54,880 --> 00:18:56,798
very basic here to begin with. If you

489
00:18:56,798 --> 00:18:58,880
think about and I love the uh

490
00:18:58,880 --> 00:19:00,640
introduction right from of all the

491
00:19:00,640 --> 00:19:02,160
background literature um which I thought

492
00:19:02,160 --> 00:19:03,119
was so cool I think is very

493
00:19:03,119 --> 00:19:04,798
complimentary the how I think about this

494
00:19:04,798 --> 00:19:06,480
as well um but if you think about just

495
00:19:06,480 --> 00:19:09,359
the raw LLM itself it's actually just

496
00:19:09,359 --> 00:19:11,919
this sequential processor that you know

497
00:19:11,919 --> 00:19:14,480
has some fixed weights and has some

498
00:19:14,480 --> 00:19:16,480
visible context and it's it's taking

499
00:19:16,480 --> 00:19:17,919
some tokens in and it's putting some

500
00:19:17,919 --> 00:19:20,640
tokens out uh to make the next decision.

501
00:19:20,640 --> 00:19:24,319
Um we don't really think of an LLM that

502
00:19:24,319 --> 00:19:26,558
way these days. We have a set of files

503
00:19:26,558 --> 00:19:29,200
that it has access to. We give endless

504
00:19:29,200 --> 00:19:31,839
programs and tools for it to use. Um,

505
00:19:31,839 --> 00:19:33,759
you can even message other sessions of

506
00:19:33,759 --> 00:19:35,279
LLMs that are going on and create sub

507
00:19:35,279 --> 00:19:37,440
agents in order to do all of these very

508
00:19:37,440 --> 00:19:39,279
cool things. But, but in its basics,

509
00:19:39,279 --> 00:19:41,279
it's just tokens in, tokens out. It's a

510
00:19:41,279 --> 00:19:43,279
neural network making a prediction. The

511
00:19:43,279 --> 00:19:46,079
harness itself is the layer between the

512
00:19:46,079 --> 00:19:48,558
LLM and the world that adds things like

513
00:19:48,558 --> 00:19:51,279
this persistent state tools and compute.

514
00:19:51,279 --> 00:19:53,440
Uh, here we have a diagram of how we

515
00:19:53,440 --> 00:19:55,839
think about prime agent. um from the

516
00:19:55,839 --> 00:19:58,079
human's perspective. So you open up your

517
00:19:58,079 --> 00:19:59,919
prime agent uh on your computer just

518
00:19:59,919 --> 00:20:01,599
like you would cloud code, codec, pi,

519
00:20:01,599 --> 00:20:03,599
etc. Um and it gets you into this agents

520
00:20:03,599 --> 00:20:05,599
view and this agents view is an overview

521
00:20:05,599 --> 00:20:06,960
of all the agents that you have going on

522
00:20:06,960 --> 00:20:08,720
for all your parallel sessions with like

523
00:20:08,720 --> 00:20:10,880
a very tight like tight summary of what

524
00:20:10,880 --> 00:20:12,640
you're using and then you can hop into

525
00:20:12,640 --> 00:20:14,640
one of those and check it out. there

526
00:20:14,640 --> 00:20:16,160
you're going to be at this root session

527
00:20:16,160 --> 00:20:17,519
and this [clears throat] root session is

528
00:20:17,519 --> 00:20:20,400
your basically the project orchestrator

529
00:20:20,400 --> 00:20:22,000
over all of these different sub aents

530
00:20:22,000 --> 00:20:23,759
that it's controlling and you don't have

531
00:20:23,759 --> 00:20:26,079
to ask it to start sub agents it will

532
00:20:26,079 --> 00:20:28,400
leverage sub agents when they're useful

533
00:20:28,400 --> 00:20:30,000
um and these are all programmatically

534
00:20:30,000 --> 00:20:32,079
called because this is all based on uh

535
00:20:32,079 --> 00:20:34,240
the recursive language model principle

536
00:20:34,240 --> 00:20:37,038
where everything is in this IPython

537
00:20:37,038 --> 00:20:39,519
shell um all your tools all your

538
00:20:39,519 --> 00:20:42,159
memories all your sub aents uh and and

539
00:20:42,159 --> 00:20:44,640
we expose uh for further coordination

540
00:20:44,640 --> 00:20:46,798
these uh these messaging paradigms. So

541
00:20:46,798 --> 00:20:49,599
you can manage all of these um and then

542
00:20:49,599 --> 00:20:50,960
these agents can directly interact with

543
00:20:50,960 --> 00:20:52,240
your environment which could just be

544
00:20:52,240 --> 00:20:53,679
like the programs and files on your

545
00:20:53,679 --> 00:20:56,079
computer or running an H200 node cluster

546
00:20:56,079 --> 00:20:57,440
maybe for your auto research. Each of

547
00:20:57,440 --> 00:20:59,038
these agents are then backed by this

548
00:20:59,038 --> 00:21:01,519
persistent dam on your computer. Uh this

549
00:21:01,519 --> 00:21:03,519
is so that you know when you close uh

550
00:21:03,519 --> 00:21:05,599
your laptop or you you control C out of

551
00:21:05,599 --> 00:21:07,038
the session, it's still running in the

552
00:21:07,038 --> 00:21:08,159
background. You have to actually stop

553
00:21:08,159 --> 00:21:09,679
the session so that make sure that

554
00:21:09,679 --> 00:21:11,759
you're continuing to running. And then

555
00:21:11,759 --> 00:21:13,279
we also have these other features we

556
00:21:13,279 --> 00:21:15,359
exposed from continual harness where

557
00:21:15,359 --> 00:21:18,000
it's able to provide like live CRUD

558
00:21:18,000 --> 00:21:20,079
operations on all of the components that

559
00:21:20,079 --> 00:21:23,599
we mentioned um in order to manage its

560
00:21:23,599 --> 00:21:26,558
uh memory skills sub agents um

561
00:21:26,558 --> 00:21:28,480
persistent and and prompt its own system

562
00:21:28,480 --> 00:21:30,240
prompt persistently. The way I like to

563
00:21:30,240 --> 00:21:32,159
think about all of this context that

564
00:21:32,159 --> 00:21:34,720
we're building up is that we have this

565
00:21:34,720 --> 00:21:37,200
sort of almost like like here I have

566
00:21:37,200 --> 00:21:40,960
like L1, L2, L3. is like a cache, right?

567
00:21:40,960 --> 00:21:42,400
It's like what is the most accessible

568
00:21:42,400 --> 00:21:44,798
information that we're working with and

569
00:21:44,798 --> 00:21:47,119
at the very like fastest like readily

570
00:21:47,119 --> 00:21:48,319
available information, you know, the

571
00:21:48,319 --> 00:21:49,440
models to be able to retrieve that

572
00:21:49,440 --> 00:21:51,038
really quickly. It's the model weights.

573
00:21:51,038 --> 00:21:52,558
So, everyone always wants to get all the

574
00:21:52,558 --> 00:21:54,400
information in the model weights. Um,

575
00:21:54,400 --> 00:21:55,599
but then we said, okay, well, maybe we

576
00:21:55,599 --> 00:21:56,640
don't have all the information because

577
00:21:56,640 --> 00:21:57,919
we don't want to have to fine-tune every

578
00:21:57,919 --> 00:21:59,038
single time to update because that's

579
00:21:59,038 --> 00:22:01,599
very expensive. So, we have this uh

580
00:22:01,599 --> 00:22:03,919
active input context. So, we're using

581
00:22:03,919 --> 00:22:06,079
lots and lots of tokens on the input. we

582
00:22:06,079 --> 00:22:08,079
might have some in context examples like

583
00:22:08,079 --> 00:22:10,720
we've seen previously in order to add to

584
00:22:10,720 --> 00:22:12,720
these different capabilities but at a

585
00:22:12,720 --> 00:22:15,679
certain point we run out of context. Um

586
00:22:15,679 --> 00:22:18,720
and so the very like earliest form of

587
00:22:18,720 --> 00:22:20,159
harnesses that we've seen that are still

588
00:22:20,159 --> 00:22:21,839
used to this day even by those who say

589
00:22:21,839 --> 00:22:23,200
we want the most minimal harness

590
00:22:23,200 --> 00:22:25,279
possible is compaction because

591
00:22:25,279 --> 00:22:28,079
compaction is a very generalized tool

592
00:22:28,079 --> 00:22:31,679
for the agent to be able to uh summarize

593
00:22:31,679 --> 00:22:33,919
its own context history in order to work

594
00:22:33,919 --> 00:22:36,319
past its context length working window.

595
00:22:36,319 --> 00:22:39,038
You can think about um once we go beyond

596
00:22:39,038 --> 00:22:40,558
like what are directly like inputs and

597
00:22:40,558 --> 00:22:42,480
outputs from the the model here into

598
00:22:42,480 --> 00:22:44,319
this L2L3. You might be familiar with

599
00:22:44,319 --> 00:22:46,319
the L3 which is more of the dispatch

600
00:22:46,319 --> 00:22:47,839
state. So if you're working with a file

601
00:22:47,839 --> 00:22:50,000
system, you can read and write from uh

602
00:22:50,000 --> 00:22:52,880
main memory. Uh if we're at the L2,

603
00:22:52,880 --> 00:22:55,519
which is I could think uh at a means in

604
00:22:55,519 --> 00:22:57,519
between uh what the active context is

605
00:22:57,519 --> 00:22:59,038
and working with your file system, you

606
00:22:59,038 --> 00:23:02,000
might have a live uh ripple, which could

607
00:23:02,000 --> 00:23:04,079
just be running things directly in

608
00:23:04,079 --> 00:23:06,319
Python uh an IPython shell like you're

609
00:23:06,319 --> 00:23:08,480
in a Jupyter notebook. And all of those

610
00:23:08,480 --> 00:23:10,000
variables are saved directly in your

611
00:23:10,000 --> 00:23:11,519
RAM. your agent can then

612
00:23:11,519 --> 00:23:13,359
programmatically manipulate them and run

613
00:23:13,359 --> 00:23:14,960
all sorts of programs directly on the

614
00:23:14,960 --> 00:23:16,880
information there saving tons of tokens

615
00:23:16,880 --> 00:23:18,480
rather than putting it directly into

616
00:23:18,480 --> 00:23:20,480
context. Uh you can also create sub

617
00:23:20,480 --> 00:23:21,679
agents and it's the same thing you're

618
00:23:21,679 --> 00:23:23,440
basically saving context here because

619
00:23:23,440 --> 00:23:25,679
you can task the agent with a specific

620
00:23:25,679 --> 00:23:27,759
set of information in order to perform

621
00:23:27,759 --> 00:23:29,519
some operations and the report back at

622
00:23:29,519 --> 00:23:31,919
the end. What is interesting here is

623
00:23:31,919 --> 00:23:34,480
what we talked about compaction for the

624
00:23:34,480 --> 00:23:36,480
active context, right? You have your

625
00:23:36,480 --> 00:23:38,240
context history. This is helping to

626
00:23:38,240 --> 00:23:39,759
update it over time. So you can continue

627
00:23:39,759 --> 00:23:42,319
to leverage this. But once we go beyond

628
00:23:42,319 --> 00:23:43,839
this, we need to be thinking about how

629
00:23:43,839 --> 00:23:45,599
are we doing these update. We talked

630
00:23:45,599 --> 00:23:47,839
about CRUD. How are we do beyond just

631
00:23:47,839 --> 00:23:49,440
creating reading? How are we updating

632
00:23:49,440 --> 00:23:52,159
and deleting our context over time

633
00:23:52,159 --> 00:23:55,119
beyond our uh or the state over time

634
00:23:55,119 --> 00:23:57,759
beyond the context line. Uh I like to

635
00:23:57,759 --> 00:23:59,359
think of this at the ripple is this

636
00:23:59,359 --> 00:24:01,038
aentic garbage collection where we're

637
00:24:01,038 --> 00:24:02,558
just cleaning up the variables in our

638
00:24:02,558 --> 00:24:04,240
state as well as like what sub agents

639
00:24:04,240 --> 00:24:06,720
could be used. And then so to make sure

640
00:24:06,720 --> 00:24:08,480
that our RAM doesn't crash my computer

641
00:24:08,480 --> 00:24:10,319
laptop every day. And then on top of

642
00:24:10,319 --> 00:24:11,919
that, we have this uh notion of

643
00:24:11,919 --> 00:24:14,000
refinement where we're updating and

644
00:24:14,000 --> 00:24:16,000
deleting the skills and memories and

645
00:24:16,000 --> 00:24:17,839
prompts that are stored on your system.

646
00:24:17,839 --> 00:24:19,359
You can think of that so that way you

647
00:24:19,359 --> 00:24:21,599
don't crash your actual uh out of space

648
00:24:21,599 --> 00:24:24,558
on your hard drive as well. And so this

649
00:24:24,558 --> 00:24:26,880
very much is a here's how we express

650
00:24:26,880 --> 00:24:29,038
this thing and here's how we revise it

651
00:24:29,038 --> 00:24:30,640
over time. The other perspective that I

652
00:24:30,640 --> 00:24:32,000
really like to think about and I really

653
00:24:32,000 --> 00:24:34,079
trying to push because harnesses are

654
00:24:34,079 --> 00:24:35,519
almost going towards this like agentic

655
00:24:35,519 --> 00:24:37,679
operating system that we're creating uh

656
00:24:37,679 --> 00:24:39,839
is when I think of it um more

657
00:24:39,839 --> 00:24:42,000
metaphorically here is that when you

658
00:24:42,000 --> 00:24:44,400
look at the raw LLM it kind of looks

659
00:24:44,400 --> 00:24:45,919
more like a touring machine where you

660
00:24:45,919 --> 00:24:47,440
have this ticker tape uh and you have

661
00:24:47,440 --> 00:24:48,798
all these instructions that are going in

662
00:24:48,798 --> 00:24:50,159
and then it's performing some set of

663
00:24:50,159 --> 00:24:52,480
operations and going out. But when you

664
00:24:52,480 --> 00:24:55,038
look at a harness it's looking a lot

665
00:24:55,038 --> 00:24:57,440
more vono like a vonoyman computer.

666
00:24:57,440 --> 00:24:58,880
you're able to do these read and write

667
00:24:58,880 --> 00:25:01,679
operations on external memory and that

668
00:25:01,679 --> 00:25:03,919
makes it much more powerful and another

669
00:25:03,919 --> 00:25:05,519
class of problems than just what a

670
00:25:05,519 --> 00:25:06,880
touring machine is able to express on

671
00:25:06,880 --> 00:25:08,960
its own. And so yeah, the idea of like

672
00:25:08,960 --> 00:25:10,400
how do you build a good hardness? You

673
00:25:10,400 --> 00:25:12,640
want it to be the most expressable thing

674
00:25:12,640 --> 00:25:16,319
you can imagine. So some some like early

675
00:25:16,319 --> 00:25:18,000
harnesses before it gets into the data

676
00:25:18,000 --> 00:25:19,278
flywheel where the models can do

677
00:25:19,278 --> 00:25:22,240
themselves are very specific. Plan, act,

678
00:25:22,240 --> 00:25:25,759
critique, do these exact specific um

679
00:25:25,759 --> 00:25:28,720
steps. Well, now har uh the models are

680
00:25:28,720 --> 00:25:30,079
able to do that themselves. You can

681
00:25:30,079 --> 00:25:31,599
imagine like we we don't have like a

682
00:25:31,599 --> 00:25:33,200
react loop that we necessarily need to

683
00:25:33,200 --> 00:25:35,440
explicitly impose. The models kind of

684
00:25:35,440 --> 00:25:37,759
have natively uh figured this out. But

685
00:25:37,759 --> 00:25:39,839
what they haven't figured out is how to

686
00:25:39,839 --> 00:25:41,359
um you know they have to be able to have

687
00:25:41,359 --> 00:25:43,599
the expressibility to call compact. They

688
00:25:43,599 --> 00:25:45,440
have to be able to have a Python ripple

689
00:25:45,440 --> 00:25:47,679
so they can run programs. Um they have

690
00:25:47,679 --> 00:25:49,599
to have the ability to programmatically

691
00:25:49,599 --> 00:25:51,599
create sub agents and access state and

692
00:25:51,599 --> 00:25:53,679
have different feedback mechanisms. Th

693
00:25:53,679 --> 00:25:54,880
those are model controlled

694
00:25:54,880 --> 00:25:56,240
expressibility features and if you

695
00:25:56,240 --> 00:25:57,599
removed one of those you're actually

696
00:25:57,599 --> 00:25:59,200
removing a capability that it won't be

697
00:25:59,200 --> 00:26:01,759
able to do otherwise. The way we manage

698
00:26:01,759 --> 00:26:02,880
um and I'm sure you're all familiar with

699
00:26:02,880 --> 00:26:05,759
the RLM paper uh from my co-author Alex

700
00:26:05,759 --> 00:26:08,159
um fantastic bit work. What we do beyond

701
00:26:08,159 --> 00:26:10,400
what was in the RLM paper is we think

702
00:26:10,400 --> 00:26:12,640
about the age sub aents as these

703
00:26:12,640 --> 00:26:14,798
persistent subsessions.

704
00:26:14,798 --> 00:26:17,679
So each the parent station can create uh

705
00:26:17,679 --> 00:26:20,558
spin up a new RLM sub aent and each of

706
00:26:20,558 --> 00:26:22,319
these are then emitted. They run some

707
00:26:22,319 --> 00:26:24,720
task and then they finish and report

708
00:26:24,720 --> 00:26:27,200
back to some end state to the parent

709
00:26:27,200 --> 00:26:28,798
session. These are then idle. They're

710
00:26:28,798 --> 00:26:30,880
still working in your RAM. At any point,

711
00:26:30,880 --> 00:26:32,640
the parent session can then send a

712
00:26:32,640 --> 00:26:35,278
message to one of the sub aents to

713
00:26:35,278 --> 00:26:36,960
continue working and it has all that

714
00:26:36,960 --> 00:26:38,960
good context that you built up over time

715
00:26:38,960 --> 00:26:40,319
so that you're not missing information

716
00:26:40,319 --> 00:26:42,159
or have to reuse information that was

717
00:26:42,159 --> 00:26:44,960
already developed in a prior context.

718
00:26:44,960 --> 00:26:46,480
And then of course you know we don't

719
00:26:46,480 --> 00:26:47,839
want to use a lot of RAM. So we can move

720
00:26:47,839 --> 00:26:50,319
them offloaded uh in an inactive state

721
00:26:50,319 --> 00:26:51,759
which then can be called back at any

722
00:26:51,759 --> 00:26:53,119
time by messaging them in this

723
00:26:53,119 --> 00:26:55,200
persistent sub agent setup. I talked a

724
00:26:55,200 --> 00:26:57,119
little bit about uh continual harness uh

725
00:26:57,119 --> 00:26:59,359
which we have in a a prior paper of mine

726
00:26:59,359 --> 00:27:02,480
which talks about cuding the entire uh

727
00:27:02,480 --> 00:27:04,400
harness state. Um this is another

728
00:27:04,400 --> 00:27:05,839
feature that we want in our coding

729
00:27:05,839 --> 00:27:08,159
agents leverage all of our prior

730
00:27:08,159 --> 00:27:09,759
history. So you can imagine like some

731
00:27:09,759 --> 00:27:11,278
set of trajectories where they have some

732
00:27:11,278 --> 00:27:12,720
actions and outcomes or something

733
00:27:12,720 --> 00:27:16,319
happened um at a at each turn. And so we

734
00:27:16,319 --> 00:27:17,919
just kind of want to expose the ability

735
00:27:17,919 --> 00:27:19,440
for the agent to leverage all that

736
00:27:19,440 --> 00:27:22,240
information in order to update what the

737
00:27:22,240 --> 00:27:23,440
future harness is going to look like.

738
00:27:23,440 --> 00:27:24,798
Are we do we need to change our system

739
00:27:24,798 --> 00:27:26,880
prompt? Do we need to create some

740
00:27:26,880 --> 00:27:29,200
skills? And skills I think of as a set

741
00:27:29,200 --> 00:27:31,759
of instructions or a program in order to

742
00:27:31,759 --> 00:27:34,960
to achieve some some specific goal. uh

743
00:27:34,960 --> 00:27:36,798
memory which could just be long-term

744
00:27:36,798 --> 00:27:38,640
storage about things that are important

745
00:27:38,640 --> 00:27:40,558
as well as the sub aent specifications

746
00:27:40,558 --> 00:27:41,759
that we talked about in this very

747
00:27:41,759 --> 00:27:43,839
persistent manner. Were there uh certain

748
00:27:43,839 --> 00:27:45,839
sub aents that we want to reuse at a

749
00:27:45,839 --> 00:27:48,000
later time because the context is useful

750
00:27:48,000 --> 00:27:49,679
and just having the ability to do this

751
00:27:49,679 --> 00:27:52,159
kind of reflection or refinement um over

752
00:27:52,159 --> 00:27:55,359
time. It's very powerful for the models

753
00:27:55,359 --> 00:27:56,880
to have. They're not perfect at this

754
00:27:56,880 --> 00:27:58,720
right now, but this is one of the the

755
00:27:58,720 --> 00:28:00,240
capabilities that you want to you want

756
00:28:00,240 --> 00:28:03,038
to build your harness such that it is a

757
00:28:03,038 --> 00:28:05,119
bit better than what the current models

758
00:28:05,119 --> 00:28:06,640
are able to do. So then you can get

759
00:28:06,640 --> 00:28:08,319
those reasoning traces and use that to

760
00:28:08,319 --> 00:28:09,919
leverage your next iteration of model

761
00:28:09,919 --> 00:28:12,240
and they'll be able to handle the

762
00:28:12,240 --> 00:28:13,440
harness and be able to bootstrap

763
00:28:13,440 --> 00:28:14,558
themselves into a higher and higher

764
00:28:14,558 --> 00:28:16,240
performance. One of the coolest features

765
00:28:16,240 --> 00:28:18,880
that we have in uh Prime Agent um that

766
00:28:18,880 --> 00:28:21,119
we we've had since the beginning of when

767
00:28:21,119 --> 00:28:22,480
I was working on this, this is one of

768
00:28:22,480 --> 00:28:24,480
the first things I added um is the

769
00:28:24,480 --> 00:28:27,278
ability to message between any any two

770
00:28:27,278 --> 00:28:29,599
agents um within like some nuclear

771
00:28:29,599 --> 00:28:31,599
family setup, parents, children, uh

772
00:28:31,599 --> 00:28:34,079
siblings. Um and the reason why I did

773
00:28:34,079 --> 00:28:35,679
this is because I was I was constantly

774
00:28:35,679 --> 00:28:36,798
trying to figure out what's the best way

775
00:28:36,798 --> 00:28:38,640
to like myself to manage all of the

776
00:28:38,640 --> 00:28:40,960
agents I have doing everything for me in

777
00:28:40,960 --> 00:28:42,480
five different directions, five billion

778
00:28:42,480 --> 00:28:44,319
different directions every day. Um, and

779
00:28:44,319 --> 00:28:45,359
it would be so much better if they could

780
00:28:45,359 --> 00:28:46,720
just like share their contacts directly

781
00:28:46,720 --> 00:28:48,398
with each other and coordinate. And

782
00:28:48,398 --> 00:28:50,480
turns out that's fantastic for like

783
00:28:50,480 --> 00:28:52,319
typical software engineering and long

784
00:28:52,319 --> 00:28:54,640
horizon jobs as well. Uh, the last thing

785
00:28:54,640 --> 00:28:57,119
that we look at when it comes to how did

786
00:28:57,119 --> 00:28:59,599
we want to design our harness is we were

787
00:28:59,599 --> 00:29:01,119
really thinking about long horizon

788
00:29:01,119 --> 00:29:03,519
performance. I want to go run some jobs

789
00:29:03,519 --> 00:29:05,038
and I don't want to have to babysit my

790
00:29:05,038 --> 00:29:06,880
agents the entire time and when when I'm

791
00:29:06,880 --> 00:29:08,558
ready to come back and check in, I can

792
00:29:08,558 --> 00:29:10,159
check in with them and see what's going

793
00:29:10,159 --> 00:29:12,480
on. And this is a perspective that I

794
00:29:12,480 --> 00:29:14,960
also really lack seeing in a lot of the

795
00:29:14,960 --> 00:29:17,519
evaluations that we're looking at. Uh a

796
00:29:17,519 --> 00:29:19,839
lot of times if you run a model for not

797
00:29:19,839 --> 00:29:21,440
enough time or say, oh well the model

798
00:29:21,440 --> 00:29:22,960
stopped working after this amount of

799
00:29:22,960 --> 00:29:25,038
budgets, but then this other model kept

800
00:29:25,038 --> 00:29:26,720
working with using more budgets. Well,

801
00:29:26,720 --> 00:29:28,159
first of all, you're not even using the

802
00:29:28,159 --> 00:29:30,000
same fixed expenditure to compare the

803
00:29:30,000 --> 00:29:31,599
models. But second of all, that could

804
00:29:31,599 --> 00:29:33,200
also be hiding performance that you're

805
00:29:33,200 --> 00:29:35,440
missing. Uh the way that I look at long

806
00:29:35,440 --> 00:29:37,278
horizon performance eval is that I want

807
00:29:37,278 --> 00:29:39,519
to see what's the practical plateau. At

808
00:29:39,519 --> 00:29:42,240
what point will we only get incremental

809
00:29:42,240 --> 00:29:44,000
gains in performance as I throw more

810
00:29:44,000 --> 00:29:45,839
test time tokens at it? I have a couple

811
00:29:45,839 --> 00:29:47,519
experiments that I'm going to show after

812
00:29:47,519 --> 00:29:49,278
we've shared design philosophy here

813
00:29:49,278 --> 00:29:52,079
about how we created uh prime agent. Um

814
00:29:52,079 --> 00:29:53,200
we're going to talk a little bit about

815
00:29:53,200 --> 00:29:55,839
test time scaling and uh our our TI

816
00:29:55,839 --> 00:29:58,720
results as well as looking at um you

817
00:29:58,720 --> 00:30:00,079
know does is it actually helpful and why

818
00:30:00,079 --> 00:30:01,359
is it actually helpful for our

819
00:30:01,359 --> 00:30:03,038
information management for the ripple

820
00:30:03,038 --> 00:30:04,079
that we're working on these long

821
00:30:04,079 --> 00:30:06,480
contexts and then um when we have these

822
00:30:06,480 --> 00:30:08,558
really really long like almost ultra

823
00:30:08,558 --> 00:30:11,759
horizon long horizon uh tasks uh how do

824
00:30:11,759 --> 00:30:13,278
we sustain these like multi-day work and

825
00:30:13,278 --> 00:30:15,119
like what actually goes on when we have

826
00:30:15,119 --> 00:30:18,000
these refinements um over like these

827
00:30:18,000 --> 00:30:19,839
settings that can last like a week at a

828
00:30:19,839 --> 00:30:21,440
time or more. So, this is a result that

829
00:30:21,440 --> 00:30:23,200
you probably all seen. We actually have

830
00:30:23,200 --> 00:30:24,960
a one additional data point that we

831
00:30:24,960 --> 00:30:26,319
added here that we didn't include in our

832
00:30:26,319 --> 00:30:28,480
original result uh just to compare

833
00:30:28,480 --> 00:30:30,640
across harnesses. We solved this uh we

834
00:30:30,640 --> 00:30:31,679
went out, we're trying to figure out

835
00:30:31,679 --> 00:30:34,240
what is the the best uh eval that people

836
00:30:34,240 --> 00:30:35,440
care about these days when we're running

837
00:30:35,440 --> 00:30:37,119
our harnesses and we're like, "Oh, we

838
00:30:37,119 --> 00:30:38,558
should do RKGI." I like, "Oh, yeah.

839
00:30:38,558 --> 00:30:40,000
Yeah, I remember. I I ran some results

840
00:30:40,000 --> 00:30:42,000
with continual harness and we got 20%

841
00:30:42,000 --> 00:30:45,519
with um Gemini Flash uh or sorry, Gemini

842
00:30:45,519 --> 00:30:47,278
Pearl." Um so, I I think we can get at

843
00:30:47,278 --> 00:30:48,720
least 20%. people who think that's

844
00:30:48,720 --> 00:30:49,839
really cool that our like general

845
00:30:49,839 --> 00:30:51,839
harness that didn't even like wasn't

846
00:30:51,839 --> 00:30:53,759
even structured for ARHI did really

847
00:30:53,759 --> 00:30:55,278
well. So I went online I was like okay I

848
00:30:55,278 --> 00:30:56,398
need to find a good system prompt

849
00:30:56,398 --> 00:30:57,599
because I don't want to make sure that

850
00:30:57,599 --> 00:30:59,519
we're losing information. So I found a

851
00:30:59,519 --> 00:31:01,278
another community leaderboard called

852
00:31:01,278 --> 00:31:02,798
prolong and I just grabbed their system

853
00:31:02,798 --> 00:31:03,839
prompt and I was like okay I'm going to

854
00:31:03,839 --> 00:31:05,599
grab their system prompt forget the rest

855
00:31:05,599 --> 00:31:06,558
and I'm just going to throw this

856
00:31:06,558 --> 00:31:08,558
directly into prime agent. Uh and then I

857
00:31:08,558 --> 00:31:11,200
ran this and I was like oh my god the

858
00:31:11,200 --> 00:31:14,640
first run that I got it hit 99.9%

859
00:31:14,640 --> 00:31:15,919
and then I looked at the logs and I was

860
00:31:15,919 --> 00:31:17,119
cheating. Okay. So, I was like, "Okay, I

861
00:31:17,119 --> 00:31:18,798
got to do proper sandboxing here. Like,

862
00:31:18,798 --> 00:31:20,480
let's set this up properly." Uh, and

863
00:31:20,480 --> 00:31:21,839
then so I spent another day on this. And

864
00:31:21,839 --> 00:31:23,119
then, and then I went back and I was

865
00:31:23,119 --> 00:31:26,000
like, "Oh my god,

866
00:31:26,000 --> 00:31:28,798
I got 78% with GPT soul. Like, this is

867
00:31:28,798 --> 00:31:30,558
going to be a great result." Um, and

868
00:31:30,558 --> 00:31:31,839
then we're back. It's like, oh, let's

869
00:31:31,839 --> 00:31:33,759
let's compare a couple other ones. And

870
00:31:33,759 --> 00:31:35,759
so, it's again, we just took the prompt,

871
00:31:35,759 --> 00:31:37,440
uh, general prompt that basically says,

872
00:31:37,440 --> 00:31:40,079
uh, use a world model to solve ARC AGI

873
00:31:40,079 --> 00:31:42,000
3. Uh, here are the actions that you can

874
00:31:42,000 --> 00:31:44,240
take. um you have uh and then the

875
00:31:44,240 --> 00:31:45,519
general system prompt for prime agent

876
00:31:45,519 --> 00:31:46,720
which is like you have a ripple you can

877
00:31:46,720 --> 00:31:49,119
call sub agents uh you can use the it

878
00:31:49,119 --> 00:31:51,599
programmatically um and and so we went

879
00:31:51,599 --> 00:31:52,720
through I went through the traces and

880
00:31:52,720 --> 00:31:53,919
it's basically doing a bunch of

881
00:31:53,919 --> 00:31:56,798
different um like calls of the coding in

882
00:31:56,798 --> 00:31:58,079
order to like check out these different

883
00:31:58,079 --> 00:32:00,000
scenarios and analyzing the images and

884
00:32:00,000 --> 00:32:01,679
doing like image processing and it's a

885
00:32:01,679 --> 00:32:04,880
lot of really cool um stuff that uh it

886
00:32:04,880 --> 00:32:06,000
seems like it was doing reasonable

887
00:32:06,000 --> 00:32:08,159
reasoning while leveraging the the

888
00:32:08,159 --> 00:32:10,558
ripple that we had um as like one of the

889
00:32:10,558 --> 00:32:12,159
main things that was able to enable

890
00:32:12,159 --> 00:32:13,679
build this. Uh so I went through and I

891
00:32:13,679 --> 00:32:15,599
ran a couple other ones. We did GPT tero

892
00:32:15,599 --> 00:32:17,919
25.7% which is really cool. You can see

893
00:32:17,919 --> 00:32:20,480
that compared to like what were the um

894
00:32:20,480 --> 00:32:22,159
like the week before we did this uh open

895
00:32:22,159 --> 00:32:24,079
AAI was like the guys the harness

896
00:32:24,079 --> 00:32:24,960
matters a lot when you're doing

897
00:32:24,960 --> 00:32:27,200
evaluations. We use the responses API.

898
00:32:27,200 --> 00:32:30,960
This is the result that we got. Um and

899
00:32:30,960 --> 00:32:34,798
we we ran Terra and and got almost like

900
00:32:34,798 --> 00:32:36,960
we we didn't run to completion this one

901
00:32:36,960 --> 00:32:38,798
but we got really good results in

902
00:32:38,798 --> 00:32:41,839
comparison. And then we go um that that

903
00:32:41,839 --> 00:32:43,359
we're already achieving higher than some

904
00:32:43,359 --> 00:32:45,038
of like the GBT soul extra high which

905
00:32:45,038 --> 00:32:48,398
was crazy. And then uh we went and we

906
00:32:48,398 --> 00:32:50,960
did Opus and hit 95.5%. We're like

907
00:32:50,960 --> 00:32:53,200
that's insane. We also compared to a lot

908
00:32:53,200 --> 00:32:54,558
of the other harnesses. So some people

909
00:32:54,558 --> 00:32:55,839
ask me like did you run this with cloud

910
00:32:55,839 --> 00:32:58,558
code? Uh I did. Um unfortunately the

911
00:32:58,558 --> 00:33:00,480
results weren't very good. Um, and so

912
00:33:00,480 --> 00:33:02,480
rather than having bad results, I just

913
00:33:02,480 --> 00:33:04,480
deferred to the the original cloud code

914
00:33:04,480 --> 00:33:05,679
results and some other people have run

915
00:33:05,679 --> 00:33:07,278
it uh with similar configurations to

916
00:33:07,278 --> 00:33:08,640
prime agent and gotten much better

917
00:33:08,640 --> 00:33:10,398
results since then. Um, but what's

918
00:33:10,398 --> 00:33:12,398
interesting is that a lot of the really

919
00:33:12,398 --> 00:33:14,558
popular harnesses don't necessarily do

920
00:33:14,558 --> 00:33:16,240
well when prime agent does well. So like

921
00:33:16,240 --> 00:33:18,398
for air agent, um, we spent a lot of

922
00:33:18,398 --> 00:33:20,640
money very quickly and uh, we had to cut

923
00:33:20,640 --> 00:33:22,319
it off because I spent like $5,000

924
00:33:22,319 --> 00:33:24,640
without making much performance. Um, not

925
00:33:24,640 --> 00:33:25,839
saying this is the best they could do,

926
00:33:25,839 --> 00:33:28,398
but it cost a lot of money to do so. Uh

927
00:33:28,398 --> 00:33:31,038
so I think that the cost to performance

928
00:33:31,038 --> 00:33:32,319
uh ratio is very important and one of

929
00:33:32,319 --> 00:33:33,759
the things that does save money is being

930
00:33:33,759 --> 00:33:36,240
able to programmatically work with your

931
00:33:36,240 --> 00:33:38,880
context. Uh we ran a bunch of long

932
00:33:38,880 --> 00:33:42,000
horizon um evals as well like oolong and

933
00:33:42,000 --> 00:33:44,159
some coding uh emulator bench which is

934
00:33:44,159 --> 00:33:45,519
going to come out soon which is a

935
00:33:45,519 --> 00:33:47,759
program bench alternative and we found

936
00:33:47,759 --> 00:33:49,839
that it was mainly parody or slightly

937
00:33:49,839 --> 00:33:52,240
better than these other harnesses like

938
00:33:52,240 --> 00:33:54,159
you across different models versus doing

939
00:33:54,159 --> 00:33:56,960
like pimono cloud codecs with glm 5.2 to

940
00:33:56,960 --> 00:34:00,319
Opus 5 and 5.6 as our setting. Another

941
00:34:00,319 --> 00:34:01,599
one I thought was really cool is we have

942
00:34:01,599 --> 00:34:04,240
this like program bench alternative

943
00:34:04,240 --> 00:34:05,679
called emulator bench where we're trying

944
00:34:05,679 --> 00:34:07,919
to reproduce entire emulators of

945
00:34:07,919 --> 00:34:09,280
computer systems or in this case

946
00:34:09,280 --> 00:34:11,119
creating like a Game Boy Color and check

947
00:34:11,119 --> 00:34:13,358
that out. And we found that what's

948
00:34:13,358 --> 00:34:14,878
really interesting is because it has

949
00:34:14,878 --> 00:34:18,000
this um ripple access in the RLM, it's

950
00:34:18,000 --> 00:34:20,320
able to use these programs in order to

951
00:34:20,320 --> 00:34:22,960
kind of do these like out of experiment

952
00:34:22,960 --> 00:34:26,320
uh loop designs in order to um try

953
00:34:26,320 --> 00:34:29,119
things out in a lot more expressable and

954
00:34:29,119 --> 00:34:30,719
free way before submitting the final

955
00:34:30,719 --> 00:34:32,639
solution to the greater. Uh we also

956
00:34:32,639 --> 00:34:35,358
tried this with uh GPU kernels um and we

957
00:34:35,358 --> 00:34:37,679
got about par results uh across

958
00:34:37,679 --> 00:34:39,918
different um both soul and kimico. One

959
00:34:39,918 --> 00:34:42,000
is better, one is worse. about par um

960
00:34:42,000 --> 00:34:43,838
which so we we're not overfit to like

961
00:34:43,838 --> 00:34:47,760
any one particular um evaluation here.

962
00:34:47,760 --> 00:34:49,760
Um what's interesting for the long

963
00:34:49,760 --> 00:34:51,760
horizon stuff is we had some auto

964
00:34:51,760 --> 00:34:54,559
research uh experiments that we did with

965
00:34:54,559 --> 00:34:56,480
the nano GPT speedrun but we scaled it

966
00:34:56,480 --> 00:34:59,039
up. We said let's give it uh 8 by H200

967
00:34:59,039 --> 00:35:02,320
for um a week and see what happens. And

968
00:35:02,320 --> 00:35:04,079
you might be like, okay, prime age is

969
00:35:04,079 --> 00:35:04,960
going to do so much better, right?

970
00:35:04,960 --> 00:35:05,838
Because it's able to do all this

971
00:35:05,838 --> 00:35:07,519
programming. Uh, it's a little high

972
00:35:07,519 --> 00:35:09,920
variance. We can't attribute um any of

973
00:35:09,920 --> 00:35:11,280
the benefits to with the harness versus

974
00:35:11,280 --> 00:35:12,719
the model there because it's a very hard

975
00:35:12,719 --> 00:35:14,559
task. But what we can do is inspect a

976
00:35:14,559 --> 00:35:16,800
lot of the behavior that we've seen. And

977
00:35:16,800 --> 00:35:18,480
what's really interesting is that we're

978
00:35:18,480 --> 00:35:20,880
seeing models like deep 6v4, GLM 5.3,

979
00:35:20,880 --> 00:35:22,960
and Kim K3. Um, you can tell these were

980
00:35:22,960 --> 00:35:24,239
done a little more recently than our

981
00:35:24,239 --> 00:35:27,280
first results. Uh, and we took these and

982
00:35:27,280 --> 00:35:28,800
they were doing like what we call out of

983
00:35:28,800 --> 00:35:30,320
loop experiments. So we were trying to

984
00:35:30,320 --> 00:35:32,000
say how can I run experiments on like

985
00:35:32,000 --> 00:35:33,280
the CPU and like look at the

986
00:35:33,280 --> 00:35:35,280
parameterization and do hyperparameter

987
00:35:35,280 --> 00:35:38,400
search and analyze the data so that I

988
00:35:38,400 --> 00:35:40,079
don't have to spend like all my time

989
00:35:40,079 --> 00:35:43,119
running expensive H200 experiments uh

990
00:35:43,119 --> 00:35:44,480
because that takes the majority of the

991
00:35:44,480 --> 00:35:46,239
time. So it's it's running experiments

992
00:35:46,239 --> 00:35:48,079
that are not the main experiment in

993
00:35:48,079 --> 00:35:49,358
order to optimize them. I think that's

994
00:35:49,358 --> 00:35:51,599
really cool behavior that we're seeing

995
00:35:51,599 --> 00:35:54,559
uh as we we shape what would be what

996
00:35:54,559 --> 00:35:56,320
kind of things we need to for the

997
00:35:56,320 --> 00:35:57,838
expressability for prime agent. So you

998
00:35:57,838 --> 00:35:59,519
can use like really good auto research

999
00:35:59,519 --> 00:36:00,639
because you can imagine if it's good at

1000
00:36:00,639 --> 00:36:02,000
auto research, it'll be good with you.

1001
00:36:02,000 --> 00:36:04,000
It be even better with a human in the

1002
00:36:04,000 --> 00:36:05,599
loop to bootstrap your experiments. Uh

1003
00:36:05,599 --> 00:36:07,760
and finally, we also streamed a 7-day

1004
00:36:07,760 --> 00:36:11,519
factorial run which used a total of 633

1005
00:36:11,519 --> 00:36:16,159
agents um across uh 23 million tok

1006
00:36:16,159 --> 00:36:18,320
output tokens in order to make like

1007
00:36:18,320 --> 00:36:20,719
steady uh tech technological advancement

1008
00:36:20,719 --> 00:36:22,480
across the tech tree to continue to

1009
00:36:22,480 --> 00:36:25,358
progress over time. And here it uh one

1010
00:36:25,358 --> 00:36:27,199
of the main benefits is I can use like

1011
00:36:27,199 --> 00:36:29,280
these sub aents that can divvy up into

1012
00:36:29,280 --> 00:36:31,280
different tasks in the factory in order

1013
00:36:31,280 --> 00:36:33,760
to research and build and gather

1014
00:36:33,760 --> 00:36:35,760
resources and build the next items to

1015
00:36:35,760 --> 00:36:38,960
design the factory. Um as well as it can

1016
00:36:38,960 --> 00:36:40,400
use the refinement to leverage what

1017
00:36:40,400 --> 00:36:42,800
happened in the past in order to help in

1018
00:36:42,800 --> 00:36:44,719
the future um over these very long

1019
00:36:44,719 --> 00:36:46,400
context so it doesn't get stuck. And one

1020
00:36:46,400 --> 00:36:47,599
of the most interesting things here is

1021
00:36:47,599 --> 00:36:48,960
that it does not get stuck and it

1022
00:36:48,960 --> 00:36:50,639
continues to make technology progression

1023
00:36:50,639 --> 00:36:54,000
even at the end of our uh stage. Um this

1024
00:36:54,000 --> 00:36:55,599
is more like a Gemini plays Pokemon kind

1025
00:36:55,599 --> 00:36:57,280
of uh conclusion here. If there's one

1026
00:36:57,280 --> 00:36:59,679
thing that uh I find interesting today u

1027
00:36:59,679 --> 00:37:00,800
but like what takeaways you should

1028
00:37:00,800 --> 00:37:02,719
actually add to your own harness. Um I

1029
00:37:02,719 --> 00:37:03,679
think that you should think about

1030
00:37:03,679 --> 00:37:06,000
agentic context management. Uh you

1031
00:37:06,000 --> 00:37:07,599
should think about swarms and looking

1032
00:37:07,599 --> 00:37:10,000
into further depth RLMs and trying to

1033
00:37:10,000 --> 00:37:11,838
run standardized eval. All of the

1034
00:37:11,838 --> 00:37:13,199
results that we can they showed today

1035
00:37:13,199 --> 00:37:15,760
can be run with our uh verifiers uh

1036
00:37:15,760 --> 00:37:18,079
package that we have at Prime Inslect.

1037
00:37:18,079 --> 00:37:20,239
Um and shout out to my collaborators who

1038
00:37:20,239 --> 00:37:22,400
are fantastic and I love working with.

1039
00:37:22,400 --> 00:37:24,054
Thanks.

1040
00:37:24,054 --> 00:37:31,119
[applause]

1041
00:37:31,119 --> 00:37:33,039
>> Hey everybody. I'm super excited to talk

1042
00:37:33,039 --> 00:37:34,800
about a project um that we've been

1043
00:37:34,800 --> 00:37:37,039
working on at Stanford. Um, I've been

1044
00:37:37,039 --> 00:37:39,519
working on this with Ivanka Orion, my my

1045
00:37:39,519 --> 00:37:41,920
co-lead author, as well as our adviserss

1046
00:37:41,920 --> 00:37:46,079
Hazeni and Christopher Ray. So, personal

1047
00:37:46,079 --> 00:37:47,760
AI is everywhere, but it's mostly

1048
00:37:47,760 --> 00:37:49,920
cloudbound today. Uh, we see lots of

1049
00:37:49,920 --> 00:37:52,559
different harnesses and projects focused

1050
00:37:52,559 --> 00:37:54,960
on making daily writing, research,

1051
00:37:54,960 --> 00:37:57,199
coding, and scheduling. But projects

1052
00:37:57,199 --> 00:37:59,519
like OpenClaw and Hermes agent typically

1053
00:37:59,519 --> 00:38:02,800
rely on cloud LMS um for most of the

1054
00:38:02,800 --> 00:38:04,639
intelligence and for most of the most of

1055
00:38:04,639 --> 00:38:07,280
the queries. Um, what does this mean? It

1056
00:38:07,280 --> 00:38:09,358
means that it's pretty costly. You're

1057
00:38:09,358 --> 00:38:11,039
getting thousands and thousands of

1058
00:38:11,039 --> 00:38:13,039
dollars in API costs if you aggregate it

1059
00:38:13,039 --> 00:38:14,960
over a year. Um, it's not private.

1060
00:38:14,960 --> 00:38:17,519
You're often sending your most personal

1061
00:38:17,519 --> 00:38:20,079
um, data to LMS up in the cloud and you

1062
00:38:20,079 --> 00:38:21,280
don't necessarily know where all that

1063
00:38:21,280 --> 00:38:23,280
data is going. Um, it also requires you

1064
00:38:23,280 --> 00:38:25,039
to rent your intelligence as opposed to

1065
00:38:25,039 --> 00:38:26,880
just simply owning it out of the box.

1066
00:38:26,880 --> 00:38:28,800
And finally, it tends to consume orders

1067
00:38:28,800 --> 00:38:30,800
of magnitude more energy than just

1068
00:38:30,800 --> 00:38:33,679
running these LMS on your laptop. And so

1069
00:38:33,679 --> 00:38:35,920
the local LMS are finally good enough to

1070
00:38:35,920 --> 00:38:37,838
actually run a lot of these queries that

1071
00:38:37,838 --> 00:38:40,800
people care about. And so we see that um

1072
00:38:40,800 --> 00:38:43,119
the the current LMS of today are only 6

1073
00:38:43,119 --> 00:38:45,679
to 12 months uh behind whatever is the

1074
00:38:45,679 --> 00:38:48,079
state-of-the-art frontier models um of

1075
00:38:48,079 --> 00:38:50,960
before. So you see um LMS today such as

1076
00:38:50,960 --> 00:38:54,320
Quen 3.8 27B um that achieve roughly the

1077
00:38:54,320 --> 00:38:57,039
same performance as like Claude 4.6 Opus

1078
00:38:57,039 --> 00:38:58,320
um back in the day. So that was kind of

1079
00:38:58,320 --> 00:38:59,838
the state-of-the-art model back in

1080
00:38:59,838 --> 00:39:02,880
August 2025. Um, and that gap seems to

1081
00:39:02,880 --> 00:39:05,358
be closing uh more and more as the

1082
00:39:05,358 --> 00:39:07,199
hardware accelerators that we have um

1083
00:39:07,199 --> 00:39:09,119
for our laptops and for our workstations

1084
00:39:09,119 --> 00:39:10,880
get better and better. Uh, just this

1085
00:39:10,880 --> 00:39:12,880
week we saw a new release from Apple um

1086
00:39:12,880 --> 00:39:14,800
with the new Mac Mini. And so we're

1087
00:39:14,800 --> 00:39:16,960
seeing this renewed focus from Apple as

1088
00:39:16,960 --> 00:39:19,039
well as Nvidia to build accelerators

1089
00:39:19,039 --> 00:39:21,519
specifically for personal use cases. And

1090
00:39:21,519 --> 00:39:23,119
so with this project, we wanted to

1091
00:39:23,119 --> 00:39:25,760
explore the the main question of can we

1092
00:39:25,760 --> 00:39:27,920
build the core of a personal AI stack,

1093
00:39:27,920 --> 00:39:30,000
namely the model inference, the agent

1094
00:39:30,000 --> 00:39:32,159
execution, the memory, the learning,

1095
00:39:32,159 --> 00:39:34,079
basically the parts that are mostly

1096
00:39:34,079 --> 00:39:36,639
reliant on the cloud today entirely on

1097
00:39:36,639 --> 00:39:38,400
device while staying competitive with

1098
00:39:38,400 --> 00:39:40,559
these cloudonly stacks. And so we

1099
00:39:40,559 --> 00:39:43,039
decided to propose open Jarvis. Um name

1100
00:39:43,039 --> 00:39:45,599
needs no needs no explanation. Um but we

1101
00:39:45,599 --> 00:39:46,960
wanted to explore just how much of this

1102
00:39:46,960 --> 00:39:48,960
we could run on device completely for

1103
00:39:48,960 --> 00:39:51,358
for free uh while preserving uh

1104
00:39:51,358 --> 00:39:54,079
security, privacy and quality. And so to

1105
00:39:54,079 --> 00:39:55,519
construct open Jarvis we wanted to

1106
00:39:55,519 --> 00:39:57,119
create the simplest set of primitives

1107
00:39:57,119 --> 00:39:59,519
for which you define any sort of harness

1108
00:39:59,519 --> 00:40:01,679
or or personal AI stack. Um the first

1109
00:40:01,679 --> 00:40:03,358
one is whatever user interfaces you need

1110
00:40:03,358 --> 00:40:05,440
to use. Um the second one is the actual

1111
00:40:05,440 --> 00:40:07,519
agentic logic around composable

1112
00:40:07,519 --> 00:40:09,280
reasoning and using different kinds of

1113
00:40:09,280 --> 00:40:11,280
intelligence and tools. Uh for the

1114
00:40:11,280 --> 00:40:12,880
intelligence, it's whatever LM you're

1115
00:40:12,880 --> 00:40:14,239
using as your engine for keeping

1116
00:40:14,239 --> 00:40:16,000
everything going. Um so this could be

1117
00:40:16,000 --> 00:40:19,838
Quen, GBDO, OSS, Gemma 3N. Um and then

1118
00:40:19,838 --> 00:40:21,599
whatever actual inference engine you

1119
00:40:21,599 --> 00:40:23,760
need to run it. So this could be O Lama,

1120
00:40:23,760 --> 00:40:27,920
um Llama CBP, VLM, SG Lang, um including

1121
00:40:27,920 --> 00:40:29,280
whatever hardware you're running it on.

1122
00:40:29,280 --> 00:40:31,280
So this could be Apple Silicon, Nvidia,

1123
00:40:31,280 --> 00:40:33,199
whatever you need. um for actually

1124
00:40:33,199 --> 00:40:34,480
making all of these agents and

1125
00:40:34,480 --> 00:40:36,559
intelligence useful you need some set of

1126
00:40:36,559 --> 00:40:38,320
tools in memory that can be run through

1127
00:40:38,320 --> 00:40:40,960
a standard MCP protocol um and you need

1128
00:40:40,960 --> 00:40:42,880
some sort of uh set of primitives for

1129
00:40:42,880 --> 00:40:44,639
actually doing learning whether it's

1130
00:40:44,639 --> 00:40:46,239
prompt based techniques like Japa or

1131
00:40:46,239 --> 00:40:48,159
DSPI um whether it's weight based

1132
00:40:48,159 --> 00:40:51,280
techniques like gpo and sftt and Laura

1133
00:40:51,280 --> 00:40:52,559
um you need some way to actually get

1134
00:40:52,559 --> 00:40:54,400
this agent to improve over time and

1135
00:40:54,400 --> 00:40:55,599
actually be able to make it more

1136
00:40:55,599 --> 00:40:57,599
personal and more effective and so to

1137
00:40:57,599 --> 00:40:59,440
kind of walk through like what opens

1138
00:40:59,440 --> 00:41:01,760
looks like um we tried to go with all of

1139
00:41:01,760 --> 00:41:04,318
the standard um interfaces that people

1140
00:41:04,318 --> 00:41:06,000
are already accustomed to. Um so we

1141
00:41:06,000 --> 00:41:07,358
wanted to give people the ability to

1142
00:41:07,358 --> 00:41:09,199
interact with it through a desktop and

1143
00:41:09,199 --> 00:41:10,480
actually just run it as they would

1144
00:41:10,480 --> 00:41:12,400
normally expect, but then see all of the

1145
00:41:12,400 --> 00:41:13,760
savings that they're getting in terms of

1146
00:41:13,760 --> 00:41:16,480
dollars and energy. Um we also wanted to

1147
00:41:16,480 --> 00:41:18,880
give people the ability um to run

1148
00:41:18,880 --> 00:41:21,199
different kinds of continuous agents. So

1149
00:41:21,199 --> 00:41:22,719
different kinds of agents that are

1150
00:41:22,719 --> 00:41:24,639
persistent in terms of cron jobs and

1151
00:41:24,639 --> 00:41:26,639
being able to run standard protocols um

1152
00:41:26,639 --> 00:41:28,880
day after day. Um, basically we just

1153
00:41:28,880 --> 00:41:30,239
wanted to to make this like

1154
00:41:30,239 --> 00:41:32,159
plug-and-play with all of the workflows

1155
00:41:32,159 --> 00:41:33,760
that people are already accustomed to

1156
00:41:33,760 --> 00:41:35,760
running. Um, and we wanted to make this

1157
00:41:35,760 --> 00:41:37,760
something that can get people to have

1158
00:41:37,760 --> 00:41:39,760
their first experience with LMS on

1159
00:41:39,760 --> 00:41:41,358
device the same way people had their

1160
00:41:41,358 --> 00:41:43,679
first experience with ChatGpt or Claude

1161
00:41:43,679 --> 00:41:46,400
back in the day. Um, so yeah, and so

1162
00:41:46,400 --> 00:41:47,760
yeah, to step through a little quicker,

1163
00:41:47,760 --> 00:41:49,519
but yeah, here's like a nice way to like

1164
00:41:49,519 --> 00:41:52,639
set up new persistent jobs. Um, yeah, we

1165
00:41:52,639 --> 00:41:54,000
have all of these different components.

1166
00:41:54,000 --> 00:41:55,519
We need some way to actually optimize

1167
00:41:55,519 --> 00:41:57,358
it. And so we wanted to get out of the

1168
00:41:57,358 --> 00:41:59,760
way of the LM as much as possible by

1169
00:41:59,760 --> 00:42:01,838
just creating a simple spec of these

1170
00:42:01,838 --> 00:42:03,280
five primitives by which they could go

1171
00:42:03,280 --> 00:42:04,880
through the optimization. And what we

1172
00:42:04,880 --> 00:42:06,318
found is that by going through this

1173
00:42:06,318 --> 00:42:08,079
whole optimization loop, we were not

1174
00:42:08,079 --> 00:42:09,519
only able to get significant dollar

1175
00:42:09,519 --> 00:42:11,838
costs uh dollar cost reductions, but

1176
00:42:11,838 --> 00:42:14,079
also significant latency reduction and

1177
00:42:14,079 --> 00:42:16,559
significant um improvements to overall

1178
00:42:16,559 --> 00:42:18,719
quality on these tests. And so this this

1179
00:42:18,719 --> 00:42:20,480
configuration is meant to simplify down

1180
00:42:20,480 --> 00:42:22,079
to just the five main things that people

1181
00:42:22,079 --> 00:42:23,440
care about when they're building these

1182
00:42:23,440 --> 00:42:25,679
these LM uh harnesses. So the

1183
00:42:25,679 --> 00:42:27,440
intelligence, the engine, the actual

1184
00:42:27,440 --> 00:42:29,920
agentic logic around it, the tools or

1185
00:42:29,920 --> 00:42:31,920
learning systems required for running it

1186
00:42:31,920 --> 00:42:34,239
um and the whole optimization um for the

1187
00:42:34,239 --> 00:42:35,920
whole spec as a whole. And so something

1188
00:42:35,920 --> 00:42:37,199
that we thought could be interesting to

1189
00:42:37,199 --> 00:42:39,280
help bridge this gap between local and

1190
00:42:39,280 --> 00:42:41,599
cloud LMS is to actually have the cloud

1191
00:42:41,599 --> 00:42:43,920
LM go through and manual and uh and

1192
00:42:43,920 --> 00:42:46,400
automatically optimize the whole LM the

1193
00:42:46,400 --> 00:42:48,159
whole local stack. And so this is a nice

1194
00:42:48,159 --> 00:42:49,519
way of taking advantages of the

1195
00:42:49,519 --> 00:42:52,318
capabilities of cloud LM to diagnose

1196
00:42:52,318 --> 00:42:55,039
proposed changes and gate um to create

1197
00:42:55,039 --> 00:42:57,440
improved solutions for these local LMS

1198
00:42:57,440 --> 00:42:59,358
while not incurring the cost of those

1199
00:42:59,358 --> 00:43:01,920
cloud LMS when you actually deploy these

1200
00:43:01,920 --> 00:43:04,159
um local stacks at inference. And so

1201
00:43:04,159 --> 00:43:06,318
what we found is that these uh these

1202
00:43:06,318 --> 00:43:08,318
open Jarvis jobs that were these open

1203
00:43:08,318 --> 00:43:10,000
Jarvis um configurations that were

1204
00:43:10,000 --> 00:43:12,480
actually optimized by cloud LMS like

1205
00:43:12,480 --> 00:43:15,358
cloud or chat GPT um were much more

1206
00:43:15,358 --> 00:43:17,838
effective than uh local stacks that were

1207
00:43:17,838 --> 00:43:19,679
just deployed out of the box because you

1208
00:43:19,679 --> 00:43:21,599
could actually cater to the specific LMS

1209
00:43:21,599 --> 00:43:24,000
the specific harness uh that was needed

1210
00:43:24,000 --> 00:43:26,000
uh for for different kinds of workloads.

1211
00:43:26,000 --> 00:43:27,838
And what we found is that even with the

1212
00:43:27,838 --> 00:43:30,159
ondevice LMS of today, we can rival

1213
00:43:30,159 --> 00:43:32,239
cloud LMS on different workflows around

1214
00:43:32,239 --> 00:43:35,119
personal AI um personal use cases,

1215
00:43:35,119 --> 00:43:37,599
coding, agentic tasks. While there

1216
00:43:37,599 --> 00:43:40,159
remains like many tasks for which um

1217
00:43:40,159 --> 00:43:42,880
like local local size LMS are not enough

1218
00:43:42,880 --> 00:43:45,440
um the gap is surprisingly closing um

1219
00:43:45,440 --> 00:43:48,239
month after month um as these LMS become

1220
00:43:48,239 --> 00:43:50,719
better distilled, more effective and

1221
00:43:50,719 --> 00:43:52,960
also we get better accelerators um for

1222
00:43:52,960 --> 00:43:54,880
running them. And so even with LM of

1223
00:43:54,880 --> 00:43:57,838
today, we can get 800x lower lower

1224
00:43:57,838 --> 00:44:00,159
costum in terms of actually running them

1225
00:44:00,159 --> 00:44:01,838
as well as a significant reduction in

1226
00:44:01,838 --> 00:44:03,838
latency. Um what we also found is that

1227
00:44:03,838 --> 00:44:05,838
no matter which cloud LM that we cloud

1228
00:44:05,838 --> 00:44:08,318
LM we picked um it was useful um in

1229
00:44:08,318 --> 00:44:10,079
terms of optimizing the whole aentic

1230
00:44:10,079 --> 00:44:13,199
loop for these uh local local open

1231
00:44:13,199 --> 00:44:15,599
Jarvis configurations. Um we found that

1232
00:44:15,599 --> 00:44:18,318
um the Opus series, Opus 5 as well as

1233
00:44:18,318 --> 00:44:20,400
GBD 5.6 Soul were were naturally the

1234
00:44:20,400 --> 00:44:22,159
best. Um but was interesting to see is

1235
00:44:22,159 --> 00:44:23,599
that you could pick Gemini, you could

1236
00:44:23,599 --> 00:44:27,838
pick um other other uh larger um LM

1237
00:44:27,838 --> 00:44:30,239
families like Kimmy and GLM and use them

1238
00:44:30,239 --> 00:44:32,400
to optimize these local configurations

1239
00:44:32,400 --> 00:44:33,679
so that you could capture those

1240
00:44:33,679 --> 00:44:34,960
efficiency gains, capture those

1241
00:44:34,960 --> 00:44:37,119
performance gains um for local inference

1242
00:44:37,119 --> 00:44:38,800
later. What we also found is that the

1243
00:44:38,800 --> 00:44:40,960
whole open Jarvis harness was cheaper to

1244
00:44:40,960 --> 00:44:43,280
optimize than alternatives which might

1245
00:44:43,280 --> 00:44:45,039
might require more data or more LM

1246
00:44:45,039 --> 00:44:47,119
calls. Um we found that like this set of

1247
00:44:47,119 --> 00:44:49,519
specs um and this set of primitives um

1248
00:44:49,519 --> 00:44:51,679
was most effective for local LM settings

1249
00:44:51,679 --> 00:44:54,960
because it got um the whole optimization

1250
00:44:54,960 --> 00:44:57,119
loop and the whole um set of LM

1251
00:44:57,119 --> 00:44:59,199
abstractions out of the way of the cloud

1252
00:44:59,199 --> 00:45:01,280
LM to just optimize the whole system and

1253
00:45:01,280 --> 00:45:02,960
just make it make it really fast and

1254
00:45:02,960 --> 00:45:04,480
really effective. Uh looking forward,

1255
00:45:04,480 --> 00:45:05,760
we're excited to keep building out this

1256
00:45:05,760 --> 00:45:07,679
project. Uh we think in the very near

1257
00:45:07,679 --> 00:45:10,159
future you're going to see um a huge maj

1258
00:45:10,159 --> 00:45:12,639
a huge proportion maybe even a majority

1259
00:45:12,639 --> 00:45:15,039
of uh people's daily inference calls

1260
00:45:15,039 --> 00:45:17,599
going to local devices and on-prem uh

1261
00:45:17,599 --> 00:45:19,920
laptops or on-prem workstations as

1262
00:45:19,920 --> 00:45:21,280
opposed to the kind of standard of today

1263
00:45:21,280 --> 00:45:22,480
where everything's being pushed out to

1264
00:45:22,480 --> 00:45:24,318
the cloud. Um we think these trends are

1265
00:45:24,318 --> 00:45:25,358
only going to continue because the

1266
00:45:25,358 --> 00:45:26,800
accelerators keep getting better and the

1267
00:45:26,800 --> 00:45:28,719
LMS keep getting better. Um and so if

1268
00:45:28,719 --> 00:45:30,719
you're excited about um anything in the

1269
00:45:30,719 --> 00:45:33,199
stack, whether it's uh better local LMS,

1270
00:45:33,199 --> 00:45:34,800
better accelerators, better inference

1271
00:45:34,800 --> 00:45:37,199
engines uh for deploying uh beyond data

1272
00:45:37,199 --> 00:45:39,440
centers, um please reach out. Uh we'd be

1273
00:45:39,440 --> 00:45:41,440
excited to chat. Include a QR code of

1274
00:45:41,440 --> 00:45:43,599
the project. Um if if folks are around

1275
00:45:43,599 --> 00:45:45,039
here afterwards, we'd love to chat.

1276
00:45:45,039 --> 00:45:52,159
Thanks. [applause]

1277
00:45:52,159 --> 00:46:00,159
have Josh and Rean.

1278
00:46:01,679 --> 00:46:03,920
we're working on QM, which is YC's

1279
00:46:03,920 --> 00:46:05,599
open-source

1280
00:46:05,599 --> 00:46:08,318
uh agent harness for work. QM is one

1281
00:46:08,318 --> 00:46:10,800
system that uh gives every employee at

1282
00:46:10,800 --> 00:46:14,318
YC an open claw-like assistant uh that's

1283
00:46:14,318 --> 00:46:17,199
like fully customizable and available in

1284
00:46:17,199 --> 00:46:20,159
Slack or via web UI, which is uh what

1285
00:46:20,159 --> 00:46:23,119
we're looking at here. Um each person

1286
00:46:23,119 --> 00:46:26,400
works within QM in their own personal

1287
00:46:26,400 --> 00:46:30,480
context that has its own sandbox files

1288
00:46:30,480 --> 00:46:34,239
uh and crons and it can they can also

1289
00:46:34,239 --> 00:46:36,400
work with QM in a multiplayer setting

1290
00:46:36,400 --> 00:46:39,760
like a slack channel. People use QM for

1291
00:46:39,760 --> 00:46:43,599
a pretty broad range of things uh like a

1292
00:46:43,599 --> 00:46:45,679
lot of automations like email triage

1293
00:46:45,679 --> 00:46:48,159
like legal and finance workflows. It's

1294
00:46:48,159 --> 00:46:49,679
really good at editing documents and

1295
00:46:49,679 --> 00:46:51,358
pulling data out of our internal

1296
00:46:51,358 --> 00:46:54,079
database. Uh it can also spin up live

1297
00:46:54,079 --> 00:46:57,358
internal web apps and help with stuff

1298
00:46:57,358 --> 00:46:59,920
like planning events. Uh but it's

1299
00:46:59,920 --> 00:47:02,559
designed to be broadly helpful for the

1300
00:47:02,559 --> 00:47:04,719
range of tasks that someone might

1301
00:47:04,719 --> 00:47:07,920
encounter at YC uh on a day-to-day

1302
00:47:07,920 --> 00:47:09,760
basis.

1303
00:47:09,760 --> 00:47:13,199
So you might be wondering uh why we

1304
00:47:13,199 --> 00:47:16,880
built this and it's really the result of

1305
00:47:16,880 --> 00:47:19,599
a string of internal agent projects that

1306
00:47:19,599 --> 00:47:21,119
have kind of unwound throughout the

1307
00:47:21,119 --> 00:47:24,400
years and all of which were really

1308
00:47:24,400 --> 00:47:26,480
riding this tailwind of increasingly

1309
00:47:26,480 --> 00:47:29,519
capable models. Um

1310
00:47:29,519 --> 00:47:32,239
the first one we built was in like

1311
00:47:32,239 --> 00:47:34,960
January of 2025. We internally refer to

1312
00:47:34,960 --> 00:47:37,119
it as the quote unquote like general

1313
00:47:37,119 --> 00:47:39,358
agent. Uh, but it was pretty

1314
00:47:39,358 --> 00:47:40,800
straightforward. Just kind of a system

1315
00:47:40,800 --> 00:47:44,480
prompt with tools uh in a loop. It was

1316
00:47:44,480 --> 00:47:46,239
one sizefits-all.

1317
00:47:46,239 --> 00:47:47,838
Um,

1318
00:47:47,838 --> 00:47:49,440
sort of like everyone was talking to the

1319
00:47:49,440 --> 00:47:51,599
same thing. Uh, and it was pretty

1320
00:47:51,599 --> 00:47:53,039
straightforward architecturally, but

1321
00:47:53,039 --> 00:47:57,519
like still uh very or surprisingly good

1322
00:47:57,519 --> 00:48:00,719
at answering data questions. Uh,

1323
00:48:00,719 --> 00:48:02,400
interestingly like the scope of what the

1324
00:48:02,400 --> 00:48:05,199
general agent was good at just increased

1325
00:48:05,199 --> 00:48:06,960
I guess unsurprisingly as the under

1326
00:48:06,960 --> 00:48:09,199
underlying models got better. Uh, and we

1327
00:48:09,199 --> 00:48:10,960
eventually hooked it up to Slack. We

1328
00:48:10,960 --> 00:48:14,960
added crons uh and gave it a a few more

1329
00:48:14,960 --> 00:48:16,800
tools so that it could be more be

1330
00:48:16,800 --> 00:48:20,800
capable across more domains. Um, in June

1331
00:48:20,800 --> 00:48:23,920
of 2025, uh,

1332
00:48:23,920 --> 00:48:25,599
by then like a lot of our engineers

1333
00:48:25,599 --> 00:48:28,159
started using cloud code and codecs. Uh,

1334
00:48:28,159 --> 00:48:29,519
and we realized that you could pretty

1335
00:48:29,519 --> 00:48:32,719
easily run these in a VM. And then, uh,

1336
00:48:32,719 --> 00:48:34,880
we hooked that up to a Slack tag, which

1337
00:48:34,880 --> 00:48:36,480
was a pretty powerful medium for people

1338
00:48:36,480 --> 00:48:38,559
who just wanted to like run a one-off

1339
00:48:38,559 --> 00:48:41,199
code change. Uh, we also configure

1340
00:48:41,199 --> 00:48:43,760
configured it to run our CI pipelines

1341
00:48:43,760 --> 00:48:46,400
and then spin up dev environments for

1342
00:48:46,400 --> 00:48:49,599
testing. Uh, and so people could come in

1343
00:48:49,599 --> 00:48:51,358
like describe a bug or something they

1344
00:48:51,358 --> 00:48:54,000
wanted to see happen and the bot would

1345
00:48:54,000 --> 00:48:55,358
go off and actually solve it, which is

1346
00:48:55,358 --> 00:48:58,400
like a pretty powerful um thing for

1347
00:48:58,400 --> 00:49:01,599
someone who like maybe hadn't uh made a

1348
00:49:01,599 --> 00:49:03,599
code change before in their life even.

1349
00:49:03,599 --> 00:49:07,280
Uh but we also on top of that had a

1350
00:49:07,280 --> 00:49:09,519
small loop going where we would observe

1351
00:49:09,519 --> 00:49:12,318
sort of how the bot failed uh where it

1352
00:49:12,318 --> 00:49:14,159
went wrong and then update the

1353
00:49:14,159 --> 00:49:17,119
agents.mmd which was uh present in the

1354
00:49:17,119 --> 00:49:20,480
codebase at the time um to make sure

1355
00:49:20,480 --> 00:49:23,440
that the thing got better as as uh we

1356
00:49:23,440 --> 00:49:26,318
like observe the usage and so in January

1357
00:49:26,318 --> 00:49:28,159
of this year a lot of the partners

1358
00:49:28,159 --> 00:49:30,639
started using openclaw and one thing to

1359
00:49:30,639 --> 00:49:32,159
know about YC partners is that they're

1360
00:49:32,159 --> 00:49:33,760
incredibly busy

1361
00:49:33,760 --> 00:49:36,480
uh between like office hours uh they get

1362
00:49:36,480 --> 00:49:38,159
tons of inbound email they're always

1363
00:49:38,159 --> 00:49:40,639
reading applications so like any tools

1364
00:49:40,639 --> 00:49:42,400
that can give them additional leverage

1365
00:49:42,400 --> 00:49:45,920
are incredibly valuable to YC so open

1366
00:49:45,920 --> 00:49:47,838
cloud particular was useful because it

1367
00:49:47,838 --> 00:49:49,280
was the first agent that a lot of them

1368
00:49:49,280 --> 00:49:51,679
had used uh that had their own computer

1369
00:49:51,679 --> 00:49:53,920
that had its own computer and so this

1370
00:49:53,920 --> 00:49:55,838
made it like very customizable in a way

1371
00:49:55,838 --> 00:49:57,920
that the previous paradigm of agents was

1372
00:49:57,920 --> 00:50:00,559
not uh and it functioned almost like a

1373
00:50:00,559 --> 00:50:02,480
personal assistant

1374
00:50:02,480 --> 00:50:04,960
And so in April uh the question became

1375
00:50:04,960 --> 00:50:06,318
like could we provide this to every

1376
00:50:06,318 --> 00:50:10,000
employee at YC uh without uh like buying

1377
00:50:10,000 --> 00:50:12,239
everyone a Mac Mini effectively. And so

1378
00:50:12,239 --> 00:50:14,639
we ended up provisioning a fleet of like

1379
00:50:14,639 --> 00:50:17,838
50 plus uh Hermes agents that were

1380
00:50:17,838 --> 00:50:20,318
running in VMs. And these were

1381
00:50:20,318 --> 00:50:22,400
definitely pretty helpful but they

1382
00:50:22,400 --> 00:50:24,480
required a lot of configuring for people

1383
00:50:24,480 --> 00:50:27,280
to get value out of them. And it was

1384
00:50:27,280 --> 00:50:28,960
just like inherently kind of difficult

1385
00:50:28,960 --> 00:50:32,000
to manage this fleet. Um, it was sort of

1386
00:50:32,000 --> 00:50:34,400
like a whack-a-ole situation where I

1387
00:50:34,400 --> 00:50:36,318
would have to sort of like SSH into

1388
00:50:36,318 --> 00:50:38,880
these individual instances and fix them.

1389
00:50:38,880 --> 00:50:40,960
And so the follow-up question became

1390
00:50:40,960 --> 00:50:43,679
like we've got we've gotten a lot of

1391
00:50:43,679 --> 00:50:46,480
value out of these agentic systems. Uh,

1392
00:50:46,480 --> 00:50:48,480
like let's build something that tries to

1393
00:50:48,480 --> 00:50:50,159
address some of the downsides of running

1394
00:50:50,159 --> 00:50:52,318
this big fleet of Hermes agents. Uh,

1395
00:50:52,318 --> 00:50:53,440
while still maintaining the

1396
00:50:53,440 --> 00:50:56,318
personalizability and some of the like

1397
00:50:56,318 --> 00:50:57,440
the stuff that people were really

1398
00:50:57,440 --> 00:51:00,318
getting value out of.

1399
00:51:00,318 --> 00:51:06,559
And so,

1400
00:51:06,559 --> 00:51:08,159
there's a pretty clear trend from

1401
00:51:08,159 --> 00:51:09,519
whether you could see there from what

1402
00:51:09,519 --> 00:51:12,159
Josh was showing you. Basically, um, the

1403
00:51:12,159 --> 00:51:14,318
models are getting better exponentially.

1404
00:51:14,318 --> 00:51:16,239
Um, and we were starting to see just

1405
00:51:16,239 --> 00:51:18,239
increasingly impressive returns from

1406
00:51:18,239 --> 00:51:19,760
giving them more and more capabilities.

1407
00:51:19,760 --> 00:51:21,838
So, OpenCloud gives the agent its own

1408
00:51:21,838 --> 00:51:23,440
computer and we start to see really

1409
00:51:23,440 --> 00:51:25,519
impressive returns from that. So around

1410
00:51:25,519 --> 00:51:27,119
May this year, we started thinking just

1411
00:51:27,119 --> 00:51:28,800
like how far can we push this if we just

1412
00:51:28,800 --> 00:51:31,119
keep pulling on this threat. Um we

1413
00:51:31,119 --> 00:51:33,119
really like the the lens of sort of

1414
00:51:33,119 --> 00:51:35,039
unhobling. I don't know if you guys uh

1415
00:51:35,039 --> 00:51:36,719
read situational awareness when it came

1416
00:51:36,719 --> 00:51:38,960
out in like 2024. Um but that was kind

1417
00:51:38,960 --> 00:51:40,800
of an era when like test time compute

1418
00:51:40,800 --> 00:51:42,480
was just starting to become a thing and

1419
00:51:42,480 --> 00:51:43,440
like you know we're starting to give

1420
00:51:43,440 --> 00:51:45,358
agents tools for the first time and

1421
00:51:45,358 --> 00:51:47,358
there's sort of this intuition that what

1422
00:51:47,358 --> 00:51:49,519
agents can do like there's a little bit

1423
00:51:49,519 --> 00:51:50,880
there's kind of more intelligence in the

1424
00:51:50,880 --> 00:51:52,719
models than we're than we're using in a

1425
00:51:52,719 --> 00:51:54,719
lot of cases. Um, and it's like if we

1426
00:51:54,719 --> 00:51:57,039
really push uh the frontier in terms of

1427
00:51:57,039 --> 00:51:58,480
just like what we're offering up the

1428
00:51:58,480 --> 00:52:00,318
agents as as capabilities that they can

1429
00:52:00,318 --> 00:52:02,719
make use of uh like magic can start

1430
00:52:02,719 --> 00:52:06,159
happening. Um so the first way that we

1431
00:52:06,159 --> 00:52:09,519
do that with QM um is by essentially

1432
00:52:09,519 --> 00:52:12,159
pulling the brain of the system up out

1433
00:52:12,159 --> 00:52:15,039
of the sandbox. So with uh you know with

1434
00:52:15,039 --> 00:52:17,838
Hermes and with OpenClaw you effectively

1435
00:52:17,838 --> 00:52:20,719
have um the agent has its own computer

1436
00:52:20,719 --> 00:52:22,239
which is super powerful but it's also

1437
00:52:22,239 --> 00:52:24,400
trapped inside that computer. So that

1438
00:52:24,400 --> 00:52:26,719
causes a few issues just from you know

1439
00:52:26,719 --> 00:52:28,318
uh us trying to administer that system

1440
00:52:28,318 --> 00:52:30,318
when there were you know uh even a few

1441
00:52:30,318 --> 00:52:31,440
dozen of these things it starts to

1442
00:52:31,440 --> 00:52:33,920
become unwieldy almost immediately. Um

1443
00:52:33,920 --> 00:52:35,760
but the other issue with that is that

1444
00:52:35,760 --> 00:52:38,880
you um all all of the like all the

1445
00:52:38,880 --> 00:52:40,800
sessions that you would have uh they're

1446
00:52:40,800 --> 00:52:42,960
trapped inside that computer. And so

1447
00:52:42,960 --> 00:52:45,838
what we did instead is we uh we just

1448
00:52:45,838 --> 00:52:47,280
offload everything into Postgress. So

1449
00:52:47,280 --> 00:52:48,719
everything is centralized um from all

1450
00:52:48,719 --> 00:52:50,239
the agent conversations that people are

1451
00:52:50,239 --> 00:52:52,480
having. Um and then we expose those to

1452
00:52:52,480 --> 00:52:54,239
the agent itself. So it can look at all

1453
00:52:54,239 --> 00:52:55,838
the context that's sort of aggregating

1454
00:52:55,838 --> 00:52:57,760
from across the system. And then the

1455
00:52:57,760 --> 00:52:59,119
other thing we do is we start thinking

1456
00:52:59,119 --> 00:53:02,000
about sandboxes less as this home where

1457
00:53:02,000 --> 00:53:04,079
the uh where the agent lives and where

1458
00:53:04,079 --> 00:53:06,400
it's kind of stuck in a lot of ways. And

1459
00:53:06,400 --> 00:53:08,400
sandboxes become more of this thing uh

1460
00:53:08,400 --> 00:53:09,838
more of a resource that the agent can

1461
00:53:09,838 --> 00:53:12,639
dip into and use as needed. Um but it's

1462
00:53:12,639 --> 00:53:15,519
it's um it's a lot less limiting. Uh the

1463
00:53:15,519 --> 00:53:17,119
other thing that this starts to open up

1464
00:53:17,119 --> 00:53:19,280
um is this idea of you're accumulating

1465
00:53:19,280 --> 00:53:20,880
this large eval set of all the traces

1466
00:53:20,880 --> 00:53:22,079
that you have from the conversations

1467
00:53:22,079 --> 00:53:23,599
that people are having with the agent.

1468
00:53:23,599 --> 00:53:25,358
And in principle, you can think about

1469
00:53:25,358 --> 00:53:27,440
going and and hill climbing on that um

1470
00:53:27,440 --> 00:53:28,639
and sort of having this automated

1471
00:53:28,639 --> 00:53:30,719
improvement loop. Um we've had sort of

1472
00:53:30,719 --> 00:53:32,400
mixed results with that. I would say I

1473
00:53:32,400 --> 00:53:34,159
think typically if you're just

1474
00:53:34,159 --> 00:53:36,000
dispatching this like torrent of agents

1475
00:53:36,000 --> 00:53:38,480
that are supposed to um fix all of the

1476
00:53:38,480 --> 00:53:40,079
bugs that they're encountering when you

1477
00:53:40,079 --> 00:53:42,159
have the LLM as a judge, you start to

1478
00:53:42,159 --> 00:53:44,159
get this kind of uh like main character

1479
00:53:44,159 --> 00:53:46,480
syndrome where the agents are making

1480
00:53:46,480 --> 00:53:48,480
fixes that are, you know, only seeing

1481
00:53:48,480 --> 00:53:51,039
their their um their piece of the

1482
00:53:51,039 --> 00:53:52,239
elephant effectively. like they're

1483
00:53:52,239 --> 00:53:55,119
they're really um they can be sort of

1484
00:53:55,119 --> 00:53:57,199
yeah not not seeing the whole um whole

1485
00:53:57,199 --> 00:53:58,639
system and so having the human in the

1486
00:53:58,639 --> 00:53:59,920
loop there has continued to be really

1487
00:53:59,920 --> 00:54:01,358
important although we're really looking

1488
00:54:01,358 --> 00:54:04,000
forward to this uh working uh all the

1489
00:54:04,000 --> 00:54:06,239
way around. Um so the other major thing

1490
00:54:06,239 --> 00:54:08,318
we do that's pretty obvious is just wire

1491
00:54:08,318 --> 00:54:09,760
the agents to all of the resources

1492
00:54:09,760 --> 00:54:11,679
across the company that we can. Uh we

1493
00:54:11,679 --> 00:54:14,000
already happened to have a CLI at YC

1494
00:54:14,000 --> 00:54:16,800
that worked really well um that wired a

1495
00:54:16,800 --> 00:54:20,000
lot of systems together. Um uh but

1496
00:54:20,000 --> 00:54:21,920
anything that wasn't in there uh we

1497
00:54:21,920 --> 00:54:25,440
basically allow um adding just arbitrary

1498
00:54:25,440 --> 00:54:28,480
API keys um that sort of thing. And then

1499
00:54:28,480 --> 00:54:30,400
we also want to ensure that we have par

1500
00:54:30,400 --> 00:54:32,159
with uh just an employee working on

1501
00:54:32,159 --> 00:54:34,239
their laptop. So you know device code

1502
00:54:34,239 --> 00:54:35,760
OOTH

1503
00:54:35,760 --> 00:54:37,838
um we go ahead and ingest that into a

1504
00:54:37,838 --> 00:54:39,199
keychain and then refresh it for you. So

1505
00:54:39,199 --> 00:54:41,679
it's it's uh ideally supposed to imitate

1506
00:54:41,679 --> 00:54:43,199
the experience of a person on their

1507
00:54:43,199 --> 00:54:46,400
computer. We mostly keep this to to be

1508
00:54:46,400 --> 00:54:48,719
read only uh in the database but we do

1509
00:54:48,719 --> 00:54:52,079
allow for rights um via uh human

1510
00:54:52,079 --> 00:54:54,159
reviewed bulk upserts. So the way that

1511
00:54:54,159 --> 00:54:55,838
works is the agent will put forward a

1512
00:54:55,838 --> 00:54:59,039
plan to um edit the database um that a

1513
00:54:59,039 --> 00:55:00,318
person can kind of give a once over and

1514
00:55:00,318 --> 00:55:01,519
ensure it's not doing anything crazy

1515
00:55:01,519 --> 00:55:04,000
before the write actually happens. Um

1516
00:55:04,000 --> 00:55:05,920
one thing we've observed with this is

1517
00:55:05,920 --> 00:55:07,679
that we've started just kind of rubber

1518
00:55:07,679 --> 00:55:09,679
stamping these. It's a little bit like

1519
00:55:09,679 --> 00:55:11,760
um I think if you guys use cloud code in

1520
00:55:11,760 --> 00:55:12,960
the early days like you might have been

1521
00:55:12,960 --> 00:55:15,119
reviewing the tool uses very closely and

1522
00:55:15,119 --> 00:55:17,440
eventually um you sort of build up more

1523
00:55:17,440 --> 00:55:18,880
trust in the agent. So this is something

1524
00:55:18,880 --> 00:55:20,480
that we're uh looking at very closely

1525
00:55:20,480 --> 00:55:22,079
over the over the next few months. Yeah.

1526
00:55:22,079 --> 00:55:24,639
So sort of like I was saying um the

1527
00:55:24,639 --> 00:55:26,800
sandboxes in this system we like to

1528
00:55:26,800 --> 00:55:29,119
think of as a resource for the agent. So

1529
00:55:29,119 --> 00:55:30,639
and and less where the agent actually

1530
00:55:30,639 --> 00:55:35,358
lives. Um so in QM the agent can um

1531
00:55:35,358 --> 00:55:37,838
basically uh by default it's going to be

1532
00:55:37,838 --> 00:55:39,679
using a particular sandbox. So that's

1533
00:55:39,679 --> 00:55:41,280
going to be one that's been allocated to

1534
00:55:41,280 --> 00:55:43,358
the user that it's talking to. But in

1535
00:55:43,358 --> 00:55:45,599
general uh there are there are

1536
00:55:45,599 --> 00:55:47,280
environments that the agent can kind of

1537
00:55:47,280 --> 00:55:49,599
converge on and uh and can collaborate

1538
00:55:49,599 --> 00:55:52,559
with. Um, and then the other uh key

1539
00:55:52,559 --> 00:55:56,079
thing is that um if the agent is working

1540
00:55:56,079 --> 00:55:58,559
on like a um like a heavier dev

1541
00:55:58,559 --> 00:55:59,838
workload, it can go and reach for a

1542
00:55:59,838 --> 00:56:01,920
machine that has more resources. Um if

1543
00:56:01,920 --> 00:56:02,960
it's working on something that's

1544
00:56:02,960 --> 00:56:04,639
simpler, it'll just go for a sandbox

1545
00:56:04,639 --> 00:56:06,719
that's that's less powerful. Um and so

1546
00:56:06,719 --> 00:56:08,159
pushing that decision into the agent

1547
00:56:08,159 --> 00:56:09,679
itself rather than the harness has been

1548
00:56:09,679 --> 00:56:12,239
a really um really powerful thing.

1549
00:56:12,239 --> 00:56:15,679
Similarly, um allowing the agent to uh

1550
00:56:15,679 --> 00:56:17,760
tap into its own runtime. So basically,

1551
00:56:17,760 --> 00:56:20,239
if it can choose the provider um that

1552
00:56:20,239 --> 00:56:22,000
it's working with, you start to get out

1553
00:56:22,000 --> 00:56:24,159
of situations like um I'm sure you guys

1554
00:56:24,159 --> 00:56:25,760
have run into this with Fable. If you

1555
00:56:25,760 --> 00:56:27,599
try to do AI research, if you try to do

1556
00:56:27,599 --> 00:56:29,760
cyber security, anything, you'll get a

1557
00:56:29,760 --> 00:56:31,599
bunch of refusals. Um so what we can do

1558
00:56:31,599 --> 00:56:32,798
when the agent can control it on

1559
00:56:32,798 --> 00:56:34,159
runtime, it can just pop out into

1560
00:56:34,159 --> 00:56:35,679
another model uh when it needs to avoid

1561
00:56:35,679 --> 00:56:39,440
a situation like that. Um similarly, uh

1562
00:56:39,440 --> 00:56:41,599
like in the earlier situation, um it's

1563
00:56:41,599 --> 00:56:43,920
often useful to pop between different

1564
00:56:43,920 --> 00:56:46,000
sandbox providers. Um and so that's

1565
00:56:46,000 --> 00:56:47,920
something we can do quite easily. And as

1566
00:56:47,920 --> 00:56:49,679
a general rule, what we've tried to do

1567
00:56:49,679 --> 00:56:52,400
is keep the harness extremely thin. Uh

1568
00:56:52,400 --> 00:56:54,000
we think of the kind of core of the

1569
00:56:54,000 --> 00:56:56,480
system as being these three tools um

1570
00:56:56,480 --> 00:56:58,400
where you have execution in a remote

1571
00:56:58,400 --> 00:57:00,480
sandbox, reading and writing from object

1572
00:57:00,480 --> 00:57:02,719
storage and then uh publishing internal

1573
00:57:02,719 --> 00:57:05,119
apps, a pretty simple sort of getbacked

1574
00:57:05,119 --> 00:57:07,519
uh system. And then we have other tools

1575
00:57:07,519 --> 00:57:09,838
for interacting with memory and crons

1576
00:57:09,838 --> 00:57:11,119
and that sort of thing. But we really

1577
00:57:11,119 --> 00:57:14,318
think of these as sort of temporary um

1578
00:57:14,318 --> 00:57:16,400
papering over rough edges in the system.

1579
00:57:16,400 --> 00:57:18,318
And really the core of it is is these

1580
00:57:18,318 --> 00:57:20,719
three up here. Um it's we we really try

1581
00:57:20,719 --> 00:57:22,480
to keep it as small as we can. Yeah,

1582
00:57:22,480 --> 00:57:23,760
we're we're sort of trying to be this

1583
00:57:23,760 --> 00:57:26,639
like uh AGI anticipating harness.

1584
00:57:26,639 --> 00:57:29,039
Although uh since we aren't there yet,

1585
00:57:29,039 --> 00:57:31,039
um there are a few things that we've run

1586
00:57:31,039 --> 00:57:34,480
into. Um, one of these is that the

1587
00:57:34,480 --> 00:57:36,239
agents have been we've tried to put them

1588
00:57:36,239 --> 00:57:37,920
in this really capable environment where

1589
00:57:37,920 --> 00:57:39,599
they have all these uh tools available

1590
00:57:39,599 --> 00:57:42,400
to them, but um they often give up way

1591
00:57:42,400 --> 00:57:44,480
too early. Uh so one thing we've

1592
00:57:44,480 --> 00:57:45,838
experimented with especially over the

1593
00:57:45,838 --> 00:57:48,239
past month or so has been setting uh we

1594
00:57:48,239 --> 00:57:50,880
call it like a grind tool or um

1595
00:57:50,880 --> 00:57:53,838
basically we set budgets on goals. So

1596
00:57:53,838 --> 00:57:55,838
the agent is not allowed to give up on

1597
00:57:55,838 --> 00:57:57,679
its task before a certain amount of like

1598
00:57:57,679 --> 00:57:59,039
walk clock time. just like a couple

1599
00:57:59,039 --> 00:58:01,039
hours uh or a certain amount of token

1600
00:58:01,039 --> 00:58:03,920
spend. And so um what that can

1601
00:58:03,920 --> 00:58:06,239
accomplish is just like uh really a lot

1602
00:58:06,239 --> 00:58:08,480
better, you know, uh research outputs,

1603
00:58:08,480 --> 00:58:10,480
better reports, that sort of thing. Um

1604
00:58:10,480 --> 00:58:12,639
and it's been really fun actually to see

1605
00:58:12,639 --> 00:58:15,199
uh like OpenAI and anthropic um you

1606
00:58:15,199 --> 00:58:16,798
know, crack some open problems in math

1607
00:58:16,798 --> 00:58:18,400
with like a very similar technique. Uh

1608
00:58:18,400 --> 00:58:20,000
but it also works for just you know,

1609
00:58:20,000 --> 00:58:22,798
normal office work stuff too. Um the

1610
00:58:22,798 --> 00:58:24,960
other thing that we've seen a lot of is

1611
00:58:24,960 --> 00:58:26,719
so this harness is supposed to work uh

1612
00:58:26,719 --> 00:58:28,159
it works in multiplayer. It works in

1613
00:58:28,159 --> 00:58:30,480
Slack. Um, but because of the the

1614
00:58:30,480 --> 00:58:32,159
artifacts of its of its training,

1615
00:58:32,159 --> 00:58:35,599
effectively uh what we see is that uh

1616
00:58:35,599 --> 00:58:37,358
the agent gets can get very confused

1617
00:58:37,358 --> 00:58:39,280
about the situation that it's in even if

1618
00:58:39,280 --> 00:58:40,798
we specify this pretty clearly in the

1619
00:58:40,798 --> 00:58:42,639
system prompt. So having like local

1620
00:58:42,639 --> 00:58:44,079
affordances for this has been something

1621
00:58:44,079 --> 00:58:45,599
that's been uh that's been really

1622
00:58:45,599 --> 00:58:53,599
important.

1623
00:58:53,838 --> 00:58:55,440
been a problem is that agents really

1624
00:58:55,440 --> 00:58:59,760
don't uh understand social contexts. Um

1625
00:58:59,760 --> 00:59:01,199
to make this a little more concrete,

1626
00:59:01,199 --> 00:59:03,199
like if I tell Regan a piece of

1627
00:59:03,199 --> 00:59:05,920
information, uh he intuitively sort of

1628
00:59:05,920 --> 00:59:08,400
knows uh or at least has like a good

1629
00:59:08,400 --> 00:59:10,559
mental framework of where it is okay to

1630
00:59:10,559 --> 00:59:13,519
share that information. Uh but it takes

1631
00:59:13,519 --> 00:59:15,838
some actual work to recreate this with

1632
00:59:15,838 --> 00:59:18,318
an agent. uh like privileged information

1633
00:59:18,318 --> 00:59:20,079
can very easily just leak into these

1634
00:59:20,079 --> 00:59:24,400
contexts where it should not be and so

1635
00:59:24,400 --> 00:59:26,000
uh the information that you can put in

1636
00:59:26,000 --> 00:59:28,000
the brain is effectively like bounded by

1637
00:59:28,000 --> 00:59:30,400
how good your permission system is uh

1638
00:59:30,400 --> 00:59:33,358
and so YC luckily has an existing

1639
00:59:33,358 --> 00:59:35,599
software system with like fine grained

1640
00:59:35,599 --> 00:59:37,599
permissioning uh that has been built

1641
00:59:37,599 --> 00:59:39,440
over the years but a lot of people just

1642
00:59:39,440 --> 00:59:42,960
don't have that uh and so it takes uh

1643
00:59:42,960 --> 00:59:45,440
work to allow for knowledge sharing in

1644
00:59:45,440 --> 00:59:48,159
an in nuanced Okay, so thanks everybody

1645
00:59:48,159 --> 00:59:51,119
for listening. Um, you can try out QM.

1646
00:59:51,119 --> 00:59:54,318
It's open source. Uh, and coding agents

1647
00:59:54,318 --> 00:59:56,079
are pretty good at standing it up. If

1648
00:59:56,079 --> 00:59:57,519
you run into any problems, feel free to

1649
00:59:57,519 --> 00:59:59,280
put up an issue and we'll look at it.

1650
00:59:59,280 --> 01:00:01,760
Uh, we're also hiring. So if any of this

1651
01:00:01,760 --> 01:00:03,519
resonated with you or would be exciting,

1652
01:00:03,519 --> 01:00:07,780
then feel free to send us an email.

1653
01:00:07,780 --> 01:00:10,780
[applause]
