Skip to main content

29th Session of the IPHC Scientific Review Board SRB029, Day 2 Part 2

Alaska News 81 min

Source

29th Session of the IPHC Scientific Review Board SRB029, Day 2 Part 2

videoAlaska News

Articles from this transcript

0:00
Speaker A

Okay, welcome back everyone. We have one final presentation here. Masha is going to give us updates on the AI approach to aging. Yeah, thank you. And just want to ask the Secretary to give me the—.

0:17
Tim

So yeah, so this is an update on using AI for self-analysis and how would each determination help And the purpose of this presentation is to summarize the current knowledge, the use of AI for determining the age of fish from images of collected bullets, and to provide an update on upcoming— try to develop an AI-based age determination. And the primary objective is to assess the viability of AI-based approaches to supplement the Civic Cloud Protection Protocol while also identifying the remaining gaps and requirements that, that potential operation should.

1:05
Basia

So in terms of key progress updates that I prepared for this meeting, first, per SMB request, I'll show the use case comparison. And this is a comparison between Break and Bake and CNN. As you see, the epigenetic aging table is in the Joseph's paper, but it's together.

1:32
Tim

I will also show the mixed method approach, which is a proposed operational design integrating manual and database aging in a single production workflow. And this is something we could discuss if we wanted to do some operational implementation. This is one thing to do. And for that particular mixed method approach, we also estimated aging error. So this would be estimation of aging error, error for 50-pixel break-and-make AI links.

2:04
Tim

As a part of other updates, I will also show the evaluation of the new model architecture. I'll focus on ConvNex model with the resolution of 480 by 480 pixels and show that it outperforms the Substrate 3 used up to pretty much last week. Now also show the assessment of performance using Z-stack metrics, which last time I was excited about—. I just got the equipment and I was building setups. So in terms of— and this is for a general—.

2:47
Tim

Just to get a little bit of a background. So reminder, the modeling approach in this project is the application of a convolutional neural network model with this type of deep learning approach. And this—. Under this approach, the layers are structured as stacks of builder is recognizing increasingly abstract features of the image. So in practical terms, the kind of the images decomposed into slices, the model recognizes the features in each slice.

3:18
Tim

So very, very different to the way single— single components find.

3:27
Tim

This is application of image regression. So we are predicting from the image. And this is because Pacific Harbor developed, as was mentioned already today, up to 55 years old, which is a little bit too much for the— tested both and the categorical predictions— and in terms of implementation, this is implementation of package TensorFlow and Keras libraries implemented in Python. And the currently preferable configuration is this ConvNeX small model. And interestingly, so the earlier testing found that alternative architectures outperform relative to Inception V3.

4:20
Basia

And as you might recall, I presented at some earlier meetings that some of them complex collapse to average. But this is— and the expectation was that this is very likely because not enough training data. And now that we have more training data, actually the more complex architectures are providing the results of that.

4:48
Tim

So that conclusion that Inception-v3 performs better as a simpler model than GloVe at the reduced size of the training pipeline.

5:05
Tim

And as a little bit of a background, since 1925, we collected over 1.6 million autoliths that have been aged and stored for potential future use. So think about it, it's a very unique resource and training tool to image them all. So huge number of boxes, boxes above the Lesum storage, are available to us, and they are already aged in traditional methods.

5:35
Basia

Aged obsoletes are sections, so they're broken in half and then baked to enhance the contrast of the corporates along the— We also have a complementary set of satellite images that are captured by geoprocessing, and those are the—. And you can compare how these look on this picture. So, this in the middle, those are the break and bake images, and on the right, the surface images. And to take those images, we use this eyepiece camera, um, it sure looks like from this the surface is just—. Is that just how—.

6:24
Basia

So they do look nicer, but you see that on the results they pretty much carry the same issue as as we see during the manual aging. So we have substantial bioavailable agents.

6:42
Tim

And when you—. So the 1.6 million liters you're taking from that collection to do this, and does this break and bake procedure then destroy that particular sample?

7:01
Basia

So, so we can't go back and take the surface. So, in order to do the surface image, it has to be complete. In order to do bake and break, it has to be both cleaned up and then we bake. So, it does not just like we can go back and do the older images of the break. And I know that, and I think we started only baking them.

7:25
Tim

So, Probably older, older oculates, they are preserved as surface images. They never—. Back in the days they were not using the break and bake method, so these we probably could do the surface. You only have a single organism per image? Yeah, so we take one image per fish unless it's a Z-stack, we'll get there.

7:51
Basia

But I mean, in the, in the stored physical atlas, there's only one per fish. Well, so for most of the fish, we have a separate collection, what we call archival collection. When there are two or more atlases, they can do left and right. These were never aged, and they are stored as unprocessed.

8:13
Speaker E

In order to age them with manual method, you need to But they can't age the, um, they only age the long side of the wings. Okay, I should have guessed there'd be some asymmetry in the name of asymmetrical fish.

8:38
Speaker A

Is that true for all flatfishes?

8:43
Tim

So we have two otoliths for some fish. Yeah, they, they are stored for potential future use, but nobody knows what that future use would be. Not yet. But they're too clean to have genetic material. That's right.

9:05
Speaker E

Earrings?

9:09
Tim

Well, I don't know, there might be some, you know, maybe in the future. I would think that they're expecting like a stroke of— who knows what would come to us. Maybe we'll hit the jackpot at some point.

9:31
Basia

So as a part of the SRB request, we put together this use case comparison. So this compares frick and frick and age-based aging. I am not think— like, I do not look at the epigenetic aging in this comparison. But if you think about the use cases, so Break&Bake is a reference aging method for different assessment, and it's a source of validated ages using available to use to train AI model. And the AI-based aging is to be used as a supplement to manual aging.

10:06
Tim

And also, we could be using a mixed method approach. So in that 50/50 mixed method approach that I'll be explaining in one of the slides. In terms of development costs, there are no— none for the break and bake. This is the established method. But if you think about this method, the real cost really is in training human nature.

10:31
Basia

Are—. These costs are considerable and it takes months to trace it. In terms of AI-based aging, the method was developed in-house as an inter— resource of intermittent staff time, and that is decreasing also. The assisted coding and the other development costs is pretty much the virtual machine cost. For testing the models and model selection.

11:01
Basia

Those models do not run on the local machines. They require more substantial GPU to— especially at this side, the training data side that we have. Pasha, how many trained agents? Currently 2, but we One of our previous agents went to grad school, so we are currently the second one is in training.

11:31
Tim

So pretty much full agent at the moment.

11:36
Basia

Probably should see another one in another course. I think that's another important use case here. So as a backup plan for also qualified staff. So, and, and, sorry, and I should say we have 2 more of our staff that are assigned aging to your food at PE. So they are trained, but they are, they know the whole process.

12:04
Basia

They can age. They are just not as efficient as a pilot.

12:13
Basia

In terms of production costs, We estimate, and it's really difficult estimation, but anything between maybe $3 to $6 per autolit, and this excludes the training, like, time training when the ager is, like, learning the whole process. And those costs are probably higher as the ager gains efficiency. So, Why would they be high? Well, because it's slow. It's just that they can do less in the same time.

12:48
Tim

As they get better, it takes them longer? No, as they get better, it takes them— so the higher as—. I mean, lower. They get lower. Yeah, yeah, yeah.

13:04
Basia

As they learn that the aging—. The aging reflex.

13:13
Tim

On the other hand, for the AI, the estimate of imaging is about 40 to 50 cents per outlet, um, and this is based on the average time it takes to take an image and the low computation needed. It takes about 15 minutes training to be able to take a search.

13:35
Basia

And the compute of the model is about $300 per model build, as you—. If you have multiple— if you have decided which model you use. So, time spent on the model selection, so just to train, in terms of benefits of break and bake, this is established method accepted for assessment and does not require any further infrastructure. On the AI side, the benefits include things like low marginal costs, and it's much easier to scale because of the time required for training. And, um, ages can be also rerun on the same images with better models.

14:25
Tim

Just use the same images. Apply better model, the model that we decided to find and just to run it on the same methods.

14:39
Basia

In terms of timelines and decision points, there are no timelines. This is established method for Breaking Faith and there's an ongoing reader position monitoring. In terms of AI-based aging, And if you think about—. I don't think we're anywhere close to thinking about full AI use, but if you think about, say, 50/50 mix, the error we estimate is actually pretty close to rate of bake only. The only decision point is we know that the temporal generalization of this model is not yet problem.

15:21
Tim

We don't have multi-generation data. That's just— yeah, so next step will be to build that, uh, image, like a training library that includes good number of years. Uh, currently we're focusing on pretty much 2019 and 2024.

15:47
Tim

But when you think about the 50/50 mix method, this offers, offers pretty much no validation change because we still age manually 50% and start training with the current year in the training. So, in terms of temporal generalization, it will be great to do that.

16:12
Tim

You think about this, all of a sudden build an application.

16:21
Tim

So, yeah, so this is the mixed method approach, and this is a proposed operational design that integrates manual and AI-based aging into single production. This is designed thinking what would be personally feasible, at the same time provides some lead to, to the agers who have a lot of bottlenecks to, to look at. The benefits include things like it's easier to scale. The CNN is updated annually. Training data accumulates also over time.

16:56
Tim

And the aging error is quantified, quantified from the quality control subset to be part of that. So the way how I thought if that was to be implemented, the best way is to split the whole, the total production collection, which is the 100% of the effluent that we designate for aging. And on the left, the green, the green box is the manual HOP, so aging 50% of production still using the traditional method, and that would feed into the AI model. So the CNN model box in the middle, which is training or fine-tuning the multi-year model. And then on the other side, we have CNN half, which is the 50—.

17:49
Basia

The other half of the production that would be derived using AI methods, but using the training component from the as well. And altogether, this 50% of the manual production and 50% of the AI would feed into full stock assessment page structure, or have 50%, yeah, composition. And on top of that, they would— and we already do that anyways, just for the current approach where we have all manual aging, there's additional QC component. So there's a 10% of autolids always aged again. So when you think about the 100% autolids, it's always 110%.

18:41
Tim

So we— that will be retained as a manual method in order to annually estimate age fitting as well, or re-estimate it to get the correct aging error for that current model and the current training data so that it's accurate for that stock assessment with this input. Inter-reader agreement, Dr. Hurst? No, it's bias-free. Yes. Because we have the break-and-make method is age validated using bomb radiocarbon, so we can assume that's unbiased.

19:20
Hurst

So when we estimate a second method, we can estimate simultaneously, estimate the bias and the precision.

19:29
Speaker A

So I think it would be sort of like a total error instead of just a precision.

19:36
Tim

Yeah, and so this is the next slide. This is about that estimation of the aging error of the, of the MIPS. So aging error was predicted for simulated 50-50 random mix of break and bake, reads, and AIC predictions. And we use that Pandadoll 2008 paper to use this joint distribution to get this aging error. And this approach fits paired age readings of the same model as Johnny estimated bias and imprecision with one method as soon as we unbiased.

20:13
Tim

With break and bake test methods.

20:20
Tim

And the results indicate that the 50/50 mix has properties very close to break and bake aging only. It's slightly less precise and slightly biased, but substantially closer to break and bake than the method we used, expanded.

20:40
Tim

And just one caveat is that the break and bake and mix estimate is based on a relatively small replicate sample size. This is based on a little bit under 1,600 bullets. So this imprecision relationship might change with additional data spanning years, but this is how many we had with those one—. 2 Reads, so we could compare really one break and bake to the second break and bake read or the mix where we randomly replaced the break and bake with the AI as a second read.

21:24
Basia

And this is the plot of these results. So if you, as you can see, we estimated the break-and-bake curve, and this is the black line, and you can see the mixed method is actually very, very close, and the band, so the confidence intervals is presented with the dashed lines. You can see that they're very close and way closer in comparison to the traditionally to the method using the path as the surface age is shown here, right? So it gets harder as they get older? Yes.

22:06
Basia

And that's why—. Why do you stop at 30? That's the maximum age you can stay. So it doesn't really matter past 30? Yeah, and I think this is just that the current sample was—.

22:21
Tim

There's not that many, so we know that the oldest The oldest we've seen in the image, the, uh, from Otto, is about 40 years old. They're not that many yet. And we're not in a targeted sample, so make sure we have those oldest. Your black blood is 100% Greek.

22:56
Speaker E

Yes. That just as an aside, the epigenetic age clock, you're going to bound 32. Uh, we have, yeah, it's 30 actually. Yeah, we have 6 from age 6 to age 30. Yeah.

23:15
Tim

So, yeah, pretty, pretty happy seeing that they're not very far, some of these, from the, the maps.

23:27
Basia

And the next slide, this is a summary of the current FIDO output overall. And here I wanted to stress that this is the same exactly set of images that I presented to bring the last SRB, but applying only new architectures, completely no change to the image input, just looking into new image, new architecture of the model. And as I mentioned earlier today, in the past, the conclusion was more complex models, they perform, they're just collapsing the average That was not—. That is no longer true now that we have a larger set of images, and now we are at the level that we can test way more, like, more wide variety.

24:24
Basia

So we tested a couple of different types, and this ConvNeX-Small model with the resolution 480 pixels performed the best. And you can see here the comparison of the results just based on the— on the—. Like, you can see here a couple of the performance metrics that show just the change based on that architecture. And you can see that in the test set, and we always send the results these are the images that the model ever seen. And currently that test set is 2,000 images out of those 11,107.

25:11
Basia

And you can see here that RMSE decreased from 168 to 162. And this is when the results are calculated for, for rounded back. So, We get the continuous value of the model, and then we want it to get the age category. And on average, the ensemble, we are using also Delsa. So we pretty much predict the same thing, different seeds, and then we get the average results from these models.

25:45
Tim

And from that, we have correctly predicted ages for 36.7% of individuals, and this is up from 35.5%, and we have additional 43.5% predictions within 1 year error, and this is up from 41%, right? And overall, this results in total agreement within 1 year for over 80% of the cases.

26:19
Tim

And I think it's important to also mention here, we—. It's hard to really expect here 100% agreement because we still don't have perfect data to train the model, but there's some degree of imprecision there based on the genetic aging error that is inherent in the data. And then the training data. So do you think that the—. Like, it looks like the AI predicted ages tend to underestimate the manual ages as the age goes up?

27:01
Basia

Yeah, so we do have a little bit of bias still using the aging methods. So this This plot is Harvard versus AI only.

27:13
Speaker B

This is where we still have considerably less training data, so that will be part of it. So you have fewer— so you can see that even in the test data, you have more younger age samples than older age samples. Is that what you mean? Yeah, so this is representing the It's just a page composition. There's always less older editions.

27:42
Basia

But like here, majority of our samples are within this 5 to 20.

27:50
Speaker B

So this is a sample of 2,000 items. Yeah, so this is the images randomly selected, 2,000 of them. For the test. Yes. But the model itself was developed on all the images you have.

28:05
Basia

It's developed on the 9,000. So, the test never goes into training.

28:14
Speaker A

Sure. Masha, one potential approach would be to set some AI-predicted age cutoff, say, AI predicts higher than that age. Then it gets read manually because you're in the area where it's starting to not— Is that logistically feasible with the workflow, or do you have to spend so much time going and finding that? So we've been wrangling with these because this is part of how you want to operationalize, right? That's what—.

28:48
Basia

On the kind of practical, yeah, you can just that, then be sure that you can— what we call fine-tune. So you can select the best predictions, just keep the predictions with the lowest CVs, and then like decide, okay, above a certain level, they do want manual aging. This is becoming problematic if you want to have the aging error going to the stock assets, but that is single estimates, then it's not so simple to get the aging. I think you can get an aging error because then you don't have—. Because it's probably going to change from year to year how many would need to go into the estimate what or how many would be designated that needs the estimation.

29:51
Tim

So it's not, it's not that clean of a method to have this feedback loop.

30:02
Speaker A

What about this during your— are you weighing boats as well?

30:08
Tim

Or fish lengths. I'm trying to—. So we do use covariates in the model. We use the geographic coordinates and we use the date when it was caught to account for the kind of where in the year. This is because we know, like, you know, if it's early in the year sample, it have a little bit less growth versus the date.

30:35
Tim

So we use the date when the fish was caught and compare it. And last time I presented the SRB, we showed that the model improved in terms of predictive power.

30:50
Tim

And we did, yeah, the geographic coordinates.

30:55
Speaker A

It's not quite what I was thinking. I'm trying to think if there's like a quick and easy way of saying whether a fish is going to be within the range of AI or if it's going to be in the upper right, then you do it manually. So we do have a method for that. So what you do, so when you do the ensemble, you have the multiple predictions. So if there are multiple predictions that are really far apart, that's a kind of indication this is not a great prediction.

31:31
Tim

So that, that this indication of actually good, good specimen to, to send for the strong verification, that, that this should be used. So because it just says that the model is unreliable, it gives you different predictions. That's the best indication that it's not live. But by that point, that botulus is off the stage and you have to go and find it, right? So I'm trying to think about something where it's real, you know, when it's, what's going up there, like, oh no, this is a, looks like it's 80-centimeter fish that's beyond the range where we AI.

32:23
Tim

You could do interference on the go. That's—. Interference is pretty cheap comparing to training. So deeper map prediction based on a previous version of the model method. Oh, so you like put it up on the stage, take an image, run the interference based on the previous one, get the prediction, and then you pops up and you're like, oh, it's manually.

32:48
Hurst

Yeah, it could be. But that's, as I said, based on discussion, we—. And those feedback loops that appear, there would be some challenges with, uh, when we fit these aging and precision relationships, we're generally assuming some kind of parametric, or at least semi-parametric, relationship between age and precision. And If we had a case where it was, say, increasing in precision to some threshold, and then all of a sudden a totally different relationship, it would be a bit of a messy model to develop. Because this actually takes a huge amount of information to estimate these.

33:27
Hurst

It's a lot because you're estimating simultaneously the probability that each method gets the same answer, but they're both wrong. Same answer that they're both right. They're off by one, but one's right. One's wrong or the other, you know, there's, there's all the combinatorics of all that. And so it actually, this is this program that Andre Punt wrote that does all of those various combinations.

33:49
Hurst

It's pretty intensive and it's pretty data hungry. And what we generally find is that you can estimate, you can estimate semi-parametric models. We start with a linear model and then we test a nonlinear model. He's got some spline options in there, but I've never seen anybody have datasets with the more complicated shapes.

34:14
Basia

So, I mean, it's not, it's not theoretically impossible, but it's— I think it might take a pretty big dataset to be able to estimate, say, a two-part relationship in the HGA precision. If we really did have a threshold age there somewhere. I guess I should have started with, this is awesome. Yeah, I mean, the performance is amazing, and I love how you thought about the workflow and how this can fit together. I just— the earlier version, there were like so many, because I was like, I wonder if we have them at each model each year, a different model, like, and then have the feedback loop, like, get the, the least precise models getting manual age and feedback for training and retraining.

35:01
Basia

And it started with a pretty complicated schematic, but then the general conclusion was we need something simpler to exactly to be able to predict those aging errors, because that was a key component. Is it getting anywhere close to something that will be acceptable? Assessment versus maybe in general you would get the higher precision, but it will be way more complicated in terms of understanding the uncertainty. Let me go back to the previous figure. I mean, this suggests to me that it's really— I mean, if you, if you can get this level of accuracy and cut your aging in half.

35:55
Basia

That's great. As you know, I've been pretty skeptical of this all along. I'm building all these additional things to try to convince in general the audience. I think that I am skeptical about going like full in, right? As a supplementary method, especially thinking about—.

36:18
Basia

I believe the challenge might be also like, this year we might have a larger survey. It's like something we have larger—. It's easy to scale versus, you know—. Yeah, I don't know. This is very amazing.

36:33
Basia

So how do you—. So the—. On the next slide, so you have a continuous response You just rounded the ages. Yeah. Yeah, and we tried the— so there are 2 ways how you can do it.

36:47
Tim

And this is based on this— in the CNN models, you have this kind of—. It's very abstract in a way, but it's like those kind of layers of the model. And then you have this last layer.

37:05
Speaker B

Again, we could retest it based on a larger sample now, but all the tests we've done—. It's interesting. So it didn't perform very well. Maybe if you have a cutout, you know, categorical variable, 20+. I'm surprised by that just because, I mean, the whole classification problem, which is like binning it into one of these groups, which is I guess effectively what a year is.

37:38
Tim

But I think this might— why the continuous one may be working better is also because there's also this agent here that's in training data. So it's like, you know, being on and off here and there, it might be just getting filled in. But there's like this error here brings Then you run— how will the age there? But in the end, we need to run because we're interested in age, like the age class. We're not interested if it's just like 2 or 6 years old, um, in which each class comes from.

38:17
Speaker A

I think the additional benefit, Tim, is that you build up this library to just As the models get better, or when I, I'm still interested in the, uh, increments, I know you said that, again, he has some initial, initial looks at that and the order of diameter is not pushable to the size across the dimensions. And that's something I was like, I just didn't have a chance to look at this more of a deep learning black box kind of approach to see if there's anything there in terms of like, if you can identify individual like, your flat band, your bullets. That's something I'm still— I'm starting to have an image library. Pretty substantial. Thank you to our interns.

39:23
Tim

Um, we also have a group of volunteers from the event that come and help us take some images. This helps a lot just getting the numbers. These spots, they seem stuck.

39:38
Speaker A

Yeah, I mean, I'm getting the, um, the improvements from the AI would be great, but most of the work is just the change of it. Yeah, having that already, even if that has to be done manually, you're already— we're doing that. Have you gone and just—. I mean, it's— you can—. There's a handful of points here that kind of look funny, right?

40:03
Basia

Oh my God. Given that there's 2,000 of them, those— like, have you looked at these outliers and and try to see the source. I just want to—. So honestly, I, I thought this might be pre-visible image, so I felt like, oh, maybe every now and then there's just somebody like, I don't know, smudge the image. There's just a really—.

40:21
Basia

Or not by accident, it's just completely not sharp. I can't visually— you can't see anything because I thought—. So there's like, you know, I run some image quality visually, I don't—. You can't see that it's seeing something. And this is actually confirmed, so what Stacks have shown, but actually the focus of the image is not important.

41:00
Basia

It's, it's, it was a big surprise. Is it a red flag? No, it's just this edge. It means that even if it's a little bit out of focus, it's, it's not problematic. It might be for us, we don't see sharp line the way how the, that these layers are filtering.

41:25
Basia

Maybe there's enough of the contrast already.

41:30
Tim

Not what we think. I know it's a sharp line for us, maybe.

41:41
Basia

So are you guys at the point where you're proposing that next year you only account—. We don't have an out-of-year sample test yet, so it's the next Yeah, so we'll get to the recommendation, but I think that the next step would be— let me go through the results. I'll just show you a little bit about also the Z-stack imaging experiment. So the Z-stack is the imaging technique in which multiple images are captured at different focal planes along the Z-axis and then combined for the single image of extended depth of field. And really the goal here was to get exactly that sharper, fully focused, auto-lit image.

42:25
Basia

And here I have a schematic of the, of the device that we built. So you can see you have a platform, and then, um, and then there's this bracket here, and this is—. To this it's attached this motorized linear stage that is holding this bracket with the little platform.

42:49
Tim

And this is how it works in practice. This is the setup, so you can see the, the, this linear stage here on the back. That was the, was the tray with the autolits. So the autolit is in this blue tray, and this is connected to a motor that is moving the the linear stage and the camera is attached to the microscope, both feed to the computer, and there's a custom software that is making those, taking those pictures automatically. And this is important because it's 80 images for each photolith at the interval of single, so it has to be automatic to make it.

43:33
Basia

So it's, but pretty much you just click and it just goes through the whole range, just image, move, image, move. And then the interns, they had the list in the software, just select the total, find out what it was. But in terms of results, so just to start with this industry workers, paired comparisons, so variants trained and tested on identical specimen setups differed only by an image input. So exactly the same specimens, but we use the inputs in the form of full stack. We also tried just using a single, a single SHARPET slice, so there's a spot that compares Which image has the highest, the largest area in the bubbles?

44:38
Speaker B

Of the 80 images.

44:42
Speaker B

And we use the previously baked and seeded image to compare, but they're exactly the same. And the only thing is, like, these burned and baked ones are sort of like The previous ones look like the ones you meant to look into. So those are surface. We use a break—. We did separate that.

45:00
Tim

So you can see here a little bit. So this is like that, like a 3D object. And that's exactly why we thought that this might be here, because there are—. They're not perfectly flat, so there's always this out of depth part of the image.

45:16
Tim

So, but the results show no accuracy. Gains, of course. So no variant beats the single sharpest slice. And the differences between the models that we tested, and there were a couple of variants of different options, it turned on and off. It's like a full range of combination of the parameters and how you try these models, but We tried 6 different variations and none of these exceeded enough.

45:52
Tim

Well, you would need like maybe 3 percentage points increase in order to be like this, to show the potential of this is really substantially better than the single. But your single sharpest slice is different from the conventional image. So you compared that too? So yeah, so we compared that as well. And actually really interesting is that the in-focus region generally shifts with depth.

46:24
Basia

So what it means is that if you just take the best slice, it does not— it's not covering the best, it's not the sharpest throughout the whole image, right? So the, the rest of the— some parts of that image are more sharp as the other slides, right? So, and this showed we run the test and 88% of the stacks we took showed that there is that variation on the— between which slides have the, the best depth of the— in the image.

47:08
Tim

But the model did not extract useful things from having available that additional sharp or more sharp parts that were available to that model from those 4-bit images. So it looks like, at least at the resolutions and of the model complexity we're currently working with, this is not providing any additional useful information.

47:40
Basia

I was hoping for more of a, wow, this is like great additional, we're going to get at least a couple percentage point improvement, but the models did not. Generally, they're pretty much So that useful operationally common means that, so if you use this, this stack, you put the ODILEC on that tray and then it automatically focuses and takes the 3D? Yes, and then so it's helpful in terms of, so currently the manual imaging that takes the, the, you know, whoever's taking that imaging to look into the microscope, they need to manually adjust the focus to get that best, at least visually. And we also compared how the single images compared to whether they are actually like as sharp as the sharpest slides. They're not, they're close, but they're not.

48:44
Speaker A

So it means just the This is not the driving force of the prediction of age. So you get a sharper, sharpest image if you just take the 80 and let it run. In theory, yes, but it's not improving that age prediction. Yeah, yeah, but I'm also thinking about in terms of the image library. So are you just going to save Sharpest slice, so you're going to save all AP?

49:14
Basia

Well, so far it was the same all AP, but this is substantially more storage required as well, big AP, which is multiple versus one. So there are definitely costs associated with this. This is— there's a storage problem with this folder because you're— what you said—. Until you crash the IPHC system to compute that the cold church space one day. We have to adjust a little bit, yeah, my software crashed.

49:52
Tim

Storage. But other than that, the storage in general is a limiting factor. You just need to, I guess, figure out how much you can what's the best way to store it outside of the Genially system. But yes, but it, uh, it removes the need for that person to choose the focal point. So it makes the focus selection automatic and repeatable.

50:27
Basia

So we think about actually implementation in the future if you want to make like an image, you know, just set up a tray with the book. I think that's a way forward.

50:40
Tim

But at the moment, there's probably not enough argument to continue with stacks compared to being saved in just images. For the input? Yeah. But operationally, you're still doing stack.

50:57
Basia

No, so it's actually stacks makes it slower because taking 80 images, it's a little slower than if one person— it's pretty simple. And you see, you don't need to actually really look into microscope because we have that software that shows you the image. Actually, maybe we could even run the sharpness test on the go.

51:21
Tim

Could you get away with taking less than 80 pictures? I mean, any of those 80 are—. So we tested a couple variants. So the model used, uh, like a band of—. So it's a Charibert slice plus minus, like, 9 on both sides, plus 9 on both sides.

51:41
Tim

Did they make a difference? So we probably could. The 80 is required because depending on how tall the OLED sets, you want to make sure you actually pick the sharpest slice. So it requires quite a—. If as an input later you don't need all 80, but if you make your base depth too narrow, you might be missing the sharpest slice.

52:08
Speaker E

OLEDs that might be sitting—. They vary by size, and then might be sitting higher. You could be increasing size of the stack.

52:18
Basia

But then I'm running into issue that I may not catch the sharpest. That's pretty, like, without that manual input increasing at like 2-3 millimeters, it's really easy to get out of focus. So if you make the intervals pretty small, you might be sharp. So I'm trying to compare what you said just now to the statement that the ZStack is still useful operationally and enables unintended batch imaging. I think you said that it's—.

52:51
Basia

So like if we build a full setup that is kind of building you like the full tray of, I don't know, 20 offlets that just takes at the time and, you know, kind of your work is to just prepare a tray and let it run, maybe then it's to be operationally efficient, but taking the autolet, putting it in, and having that like a 1 minute in between, there's not like you can do something else, and that's too short of a time to do something else in between changing the autolets, but it takes time.

53:27
Tim

I mean, are you thinking that you have another mechanical stage that's That's true. I think that's a little bit—. I think we'll have to have a little bit more arguments to really need it. I think it would require quite a more, more, more fancy, I guess, to mechanically set up, and then more thinking about the the conditions around. It will be a little bit more complicated engineering-wise.

54:04
Tim

The 35 seconds that you mentioned before, is that for the whole stack? 35— So that was for the single chip. For the stack, I think it takes at least a week or two And it's all together, I think it takes about a minute to go through the T-shirt.

54:31
Basia

But it's not only that, it takes a moment to like, it's not, you take one, it's not very long to just save an image, but think about saving 8 images. It takes not a huge time, but when you thousands of these.

54:54
Basia

I need to check. I can check exactly how. Yeah, is the technician inputting the— typing in the number or sample number? No, so it's, um, you pretty much, you pick the box and check the vessel name and you just select the the sample name and you have the list of just start and it's kind of automatically it goes to the name and you just need to skip.

55:27
Speaker A

I'm happy to share with the whole seminar. So you guys don't do barcoding samples?

55:35
Tim

No, they're in these boxes that are 100 samples. They're kind of lined up The price. I think the big challenge is like the, the cell is maybe like centimeter by centimeter. So need to be rid of that board. Okay.

55:57
Basia

There will be open spot. Having not used a microscope in a long, long time, how long does it take to bring it in focus? It's great, like if there's just like one app that you turn on, you don't even need to know how to use microscope in order to see.

56:15
Basia

And as I said, it doesn't seem to be— and I— that's what's— yeah, it might be, I assume, like, you know, at some point it might be much more of information area. So that you can maybe find them better, but just not the scale set.

56:39
Speaker B

And those images are originally taken as 2,500 by 2,500 pixels for thinking about the production feature use, so we don't decrease that permanently. But since the current model, the largest model that was Oh, sorry, it's resized to scale to the—. I just had one question about what you said. So there's the— so there's the bake and break, the manual. Then you said we kind of know there's not much error associated with this, carbonated or— so is there a set of these bake and break samples that you also have that information?

57:28
Hurst

So that was a study that was done quite a while ago now, maybe close to 20 years ago. It's using the bomb radiocarbon dating method. So they did— they actually—. That whole method was developed based on halibut, and they— so they developed the curve of the increase in that isotope. And they match that up to fish collected in certain years and the ages associated with those.

57:59
Hurst

And that's how we can validate that the break and bake method is unbiased. There's, there's imprecision, but on average it works.

58:10
Hurst

And so that, that method's now been applied to a bunch of different species. It's first done for halibut.

58:18
Hurst

So we don't, we don't have a true known age, say. But you know, that's actually compared to most species, it's amazing. Precise. Very convenient method. I first got here, I couldn't pick up good data.

58:34
Speaker E

And you can also just look at the age data. You can see cohorts going up the diagonals, and that's good.

58:41
Hurst

Species where you either can't see cohorts, maybe they don't have variability, but maybe you just can't see the cohorts, or even worse, in those cases where you see a cohort and then it skips, you know, skips to a different diagonal at its age. That's also a case.

59:00
Basia

We had some challenges in the past where exactly, like, the cohorts is not aligned. I was like, oh yeah, we need to re-age some of them.

59:11
Speaker B

The Punt paper that you refer to, which is basically an error in verticals problem for you, so it was worried about precision and bias in the— although you didn't have it. Well, you have to have— for the method to work, you have to have at least one—. For the Punt approach to work, you have to have at least 1 aging method that's unbiased, or you have to assume that it doesn't have to be perfect, but it has to be unbiased. So you can estimate the precision for N methods and you can estimate the bias for N minus 1. But you have to have some way to anchor that in.

59:50
Hurst

Sometimes people do that with known age samples, subset of just assume that.

1:00:00
Hurst

But in our case, it's actually validated.

1:00:03
Speaker E

But that's, again, that's a pretty data hungry because you're estimating all of that information.

1:00:12
Tim

2018, 2008. Yeah, and he's got a GitHub site downloaded.

1:00:20
Tim

Do it. It's not— it's, it's physically terribly intuitive, would you say? But it's okay. And we re-estimate like we —this is something too. So, like, not just take the estimates as they were estimated for the stock assessment some time ago now.

1:00:37
Speaker E

We re-estimated it to get the— based on the data we have currently, we estimated that. Our data set for replicate ages on break and bake is like 7,000. So we have some some power in the analysis. So, even though we only had 1,500 for the, for the mixed method, we have enough. And that's one of the things we wanted to re-estimate just to make sure it's lining up, which I did the same thing 10 years ago.

1:01:13
Speaker E

Then we had 61,000, you know, we still had—. So, Joseph, in your elastic net model, are you assuming that Each is perfect. We haven't got that yet. You could do the pun thing first.

1:01:33
Speaker E

250 Use, I'm certainly not going to be enough to—. That's the tricky part, sample size. I guess I think in Andre's paper they recommend at least 50 per age band. So, I don't know, that's kind of more— you're including the CPGs that are associated with age, so you don't know how many there are, but they may be in the order of 100,000. Yeah, but no, no, my point was that it's—.

1:02:05
Tim

Your response is still something that's got some noise in it because your response is an age—. Your— it's age as a function of a gazillion CPG sites. The age isn't known exactly. And so do you— I mean, the question is, do you bother? Maybe those were—.

1:02:28
Basia

But they're also using multiple, like, the ages that were coming from multiple reads too. Yeah, they're the double reads. I mean, that's the black line. It's a black line. Okay, I forgot Uri had the double read.

1:02:45
Hurst

Yeah. And I think there's a pretty important difference in the objectives there. So in that case, they're trying to get the best estimate of actual age. In this case, what we're trying to do is develop a method that's replicable, that we can define the properties of. So we actually don't want to only limit to really well-read otoliths because in the future that's not going to work.

1:03:06
Speaker B

Kind of go with the lines here, but it has to be applicable to everything that comes out of the sausage factory. Yeah, no, I understand that. I just thought it's just a— I think it's an interesting source of error that probably in your case wouldn't be too worried about.

1:03:29
Speaker E

I think pretty understood that there's in the chronological age there's be a slight error there.

1:03:36
Basia

But yeah, I mean, in our hands, that's as little error as we can get. Yeah, because otherwise, I mean, the only like known age is otherwise it's a species that they can raise differently.

1:03:51
Hurst

Yeah, like, yeah, it's like for fish, you can't—. They would grow differently. Tetras like being marked. Or tag recoveries. There's only a few ways you can get your—.

1:04:03
Mike

Mike, go ahead. Oh, I was wondering if you guys had been thinking through how you might try to test the effects of a change in aging method on the assessment results before adopting it, because it seems to me like this would actually be a pretty big change in the aging error matrix, even though you're getting remarkable, remarkably good fits and they seem to be getting better even since May, that still I'm wondering how much of the information on fishing mortality and things in the model is coming from those older ages and considering how having Yes, it gets accounted for in having an aging error matrix, but that has to add a lot of noise, and I'm just wondering what it actually does to the ability to estimate to the overall model accuracy and things like that.

1:05:09
Hurst

Yeah, we've thought a lot about that. I mean, a couple, couple things there. We already have two aging error matrices in the model. We have surface ages and break-and-bake ages. And actually, the, the imprecision matrix for this, for a 50/50 mix, as Basia showed, is pretty close to the matrix for the break-and-bake method.

1:05:35
Hurst

This is, this is the, the expectation of that matrix and the 95% confidence interval of ages. So it's to just imagine a distribution fit to the black line and the blue lines. They're pretty close. In terms of testing it, I think it's— when we started down that road, I mean, it's really easy to throw some ages in or change the imprecision matrix. But thinking about what we're actually trying to test, because if we want to see, well, if we did this for a few years, what would the effect that's completely different than if we go to a method that's much less precise and we do that for an entire generation of fish.

1:06:16
Hurst

So you're really talking about a simulation experiment that would have to include the timeline that you're intending to apply this over and whether or not you have all, all of the sources of ages coming from this or whether we apply this to say the Fishery data but we keep doing the survey manually. There's really a pretty rich suite of hypotheses that you need to think about there. Yeah, yeah, I completely agree. And I think the— I think that whole suite of options just lends itself to being informed by trying to think about them in a simulation context. The one thing I would narrow it down on is you're obviously are you planning on adding these as additional years to the current data in the model?

1:07:08
Mike

And so you're thinking about having your whole historical dataset and you're going to be adding to it in the future. And so I would be thinking in a simulation in exactly that context, not a more generalized one.

1:07:22
Hurst

So we've already done that experiment because we transitioned from surface aging to break and bake. In 2002. So we used to use— all of our assessment was based on the red lines, and then we made that transition, and then since then we've been in break and bake. So we have— we've actually done— effectively done this once, although it was to a more precise rather than a less precise method. I guess— I guess I look at the difference between the blue and the black lines here And that doesn't tell me that this would be worth doing a simulation experiment on.

1:07:58
Mike

Oh, that most places have aging methods that are far less precise even than this 50% mix method. I agree, but I think the effects of aging error are not well, well understood and well recognized. I think they're, they're more insidious than people people give them credit for.

1:08:27
Hurst

Absolutely. We actually did a paper with Kevin Peiner years ago where we looked at estimating the bias in aging and precision internal to the stock assessment model. And we showed that you can actually propagate some of this uncertainty into the management results if you do the whole thing inside the assessment model. But let me tell you, it's painful. We actually programmed it into stock synthesis and estimated the bias simultaneously with the rest of the model parameters.

1:08:58
Hurst

And it was quite an exercise. I have to admit, at the end of the day, we couldn't recommend that anybody actually do it in a production setting. But it was— we did demonstrate that you could do that and you could propagate some additional uncertainty by doing so. I think it's kind of— I think Mark Mander did a similar exercise with CPE data showing that you can do a full CPE standardization from raw data inside your stock assessment model, and you get a better translation of uncertainty to the final product, but it wasn't worth it there either. Yeah.

1:09:35
1:09:41
Tim

I guess one important thing to keep in mind here, the way how we were thinking about when designing that, like, potential operation might take it is to keep it adaptable and flexible, and the manual component there is equal to that quality control sample that is available to, to monitor later and re-estimate it every time because, for example, we don't have kind of propagate the increasing error that we don't know what the solvent was.

1:10:15
Speaker A

So why 50/50? Yeah, that was just—. I mean, that's, that's certainly where I would start, but I mean, can you go to 90/10?

1:10:27
Basia

Oh, I think we, we did not—. So you can compute, I can run it, uh, it's probably gonna be, you know, getting closer to just what I showed here. This is pretty much, uh, this one is 0, 100, right? This is probably going to be somewhere in between. Um, this was thinking more about what this kind of— what could be potentially feasible, but would make people be still comfortable with to have the, you know, still the sample, good number of samples that are still aged in traditional methods.

1:11:14
Tim

But at the same time, there's enough to, to reduce enough the number of manual aging to make a difference. In terms of how many we have to age every year. So I think it's a good starting point. I don't know what else would you use to think about optimizing, you know, optimizing performance. It goes back to what you said before.

1:11:51
Basia

But I think it will be pretty much— it, it would be a decision, right? Where, how far, like, you know, it's really decreasing that decision a little bit, how far you go.

1:12:10
Basia

Um, yeah, but in general, I just have one last slide with conclusions. So this is Just summarizing what I was saying today, and the ensemble is based on the deep learning CNN model is now human reader agreement. We have 80% within 1 year, 74.4% within manual reading competence interval. So that's a good one. This is the image we selected for this presentation.

1:12:42
Basia

The architecture matters. The newer backbone is called MAG, so we also tried addition and the 2M, which performed pretty close and now outperformed the earlier Inception V3 model that we were presenting up to the last meeting. And this is supposed to be enabled now with a larger training dataset. We show here proposed 50/50 brick-and-bake and AMX method, which gives us same total aging efforts, so 110% of the ages, same as previously, but specialist manual burden is roughly half, so 60% under this approach. Aging error is the— of the mix is close to break and bake, slightly less precise slightly biased but substantially closer to brick and bake if compared to, to service and shopping, that red line.

1:13:48
Basia

And there's a built-in adaptation to change of 50% fresh expiries every year without them trying to predict or track interannual trends about the demand. So I think adaptation. It's really built in. And then there are also a couple of support bindings. I did not mention it in the presentation, but surface imaging remains forced and break and bake.

1:14:17
Tim

And this pretty much looks like it's carrying the same issues as from Edge. So we have same level of bugs coming with the surface images as we have from surface. As these stack images give no accuracy gains in the final depth of context. In terms of next steps, uh, suggest that extending, extending training and testing across multiple years, decades, to confirm year reliability of this. So the next step, what we're thinking, is just for So have a good sample of just a couple of decades and have some, uh, policy might be very good sense to see if we're gonna smooth your— it's not— it doesn't have to be a single year model, a multi-year model, that means just multi-year, it's still predicting fine for test method and not seeing images.

1:15:27
Speaker B

But the human expertise, I think, remains precise just because how valid that method has been. I just have one—. This is probably silly question— just so I understand, because it's so interesting. So you have the break and bake rate. So all the person doing, um, knows is they have seen what year it was caught.

1:15:55
Speaker B

They don't use any of that extraneous information to age. CNN Ensemble, you said it's got all these layers that are derived from the image, but you also input the year it was collected, where it was. Not the year, because so because Most of the tests we did, we don't currently have a small variation, so we did not. There are other inputs, but we included the day of the year to kind of capture the seasonality.

1:16:27
Tim

Okay, so day of the year and the spatial component as well.

1:16:34
Speaker B

So the, the bake and break version, are they all collected at the same time of the year? Like, I'm just trying to understand why you have one. They know when it was collected. They know, and they do— they do—. I don't know how exactly it works, but they do take into account the edge.

1:16:53
Speaker B

They do? The manual agers? Okay, there's more intuition. I mean, whatever. Yeah, they, they, they—.

1:17:00
Hurst

Based on like the season versus like if they look And they know what year it was collected and then they actually know the source of it, you know, it's Fishery Survey, but they don't use the length to infer the age. There's, I'm just trying to make, so that make sure the model reflects the same information as, you know, like you want it to mimic what that ager knows and use to make their prediction. And I think that, but these things here, I think that will be potentially variable as a covariate if you have like this—. That would be a model tuning thing you could test. Yeah, and you could even have a smooth, like a seasonal effect rather than the day of year if you want to.

1:17:52
Tim

So currently it's included as a continuous variable, so it's pretty much like which day of the year. So it's like day, month is converted to a picture.

1:18:04
Basia

What should I forget? The analysis of cross-validation across regions. Yes, so we did that. Then we found that there were no significant differences if the model was trained including or excluding the region. So, we did this kind of full matrix for each individual regulatory aspect.

1:18:29
Basia

There was this matrix of like prediction, including. So, we did that. So, the conclusion was that, uh, it did not prevent the model to—. Images from the different regions. And we also think Generally, we think that this would not be that much of a problem because moving forward, we already have images from all the regions.

1:18:55
Tim

So there will be some images.

1:19:01
Tim

Any more questions for Pasha about this?

1:19:19
Speaker E

Okay, do we go to break and then we'll come back? Sounds good. And then so for tomorrow, what time do you want to convene? You also have a discussion session there, but public session, what time?

1:19:38
Speaker E

At the moment on the agenda, it's after lunch. Sorry, not after lunch.

1:19:45
Speaker E

11:00. 11:00. After break.

1:19:53
Speaker E

So for those listening, we'll reconvene tomorrow in the public session at 10:30 AM.

1:20:05
Speaker A

Anna and Mike, please stay on.

1:20:10
Speaker A

Thank you very much.

Speakers in this transcript