# Transcript

**Boris Veytsman**: [09:40] Okay, friends, could you please take your seats?

**Boris Veytsman**: [09:48] Okay, I guess we have, how many seconds before start? Oh, exactly. Okay, I am Boris Veytsman. I am outgoing Vice President of TeX Users Group. According to our rules, the new Vice President takes office at AGM, Annual General Meeting on Sunday. So I can present to you Erik, who is going to chair today's session, and he is Vice President-elect.

**Boris Veytsman**: [10:23] And as a Vice President, I have an honor to open this conference. For the first time, we are meeting in Kerala. I think it's the record, and it's great, and the previous two times were fantastic. And from what I already have seen in the workshop, in the reception yesterday, in the program, both technical program and cultural program, is going to be fantastic as well.

**Boris Veytsman**: [10:54] I am going to say that most of this is because of great work by STM Docs, who is the main sponsor of this conference, and the organizing committee. I want to thank all of you for making all this possible, including those who are now remotely, Karl and Sophia, who are, I hope, watch this through YouTube. And I also would like to thank our sponsors, besides STM, it's the users groups, Dante and TeX Users Groups and TUG. It's Elsevier, Google, Overleaf, and [Cerda?], some of the representatives of these groups are here. And thank you for your really generous sponsorship and generous help. We would not be able to do anything without you, and it's very much appreciated.

**Boris Veytsman**: [11:56] A couple of technical things. I was asked to tell you to please switch off your mobile phones or put them in the vibrating regime or whatever. We are not in theater where they would just throw you out for a buzzing phone, but still, please do. Second, for those who are remote, if you are watching us, there is a chat. You can ask the questions or comments in the chat, and if time permits, we will ask them here. And with this, I would like to give the baton to Erik, who is the chair of today's session, and Erik, please go ahead. And oh, I forgot to say, okay, officially, TeX Users Group meeting of 2023 is open.

**Erik Nijenhuis**: [13:00] I am Erik Nijenhuis. I will be your session chair for this morning. And we will start with the talk of Rob Schrauwen from the Netherlands. Rob is based in the Netherlands. He is currently a VP Data and Platform Strategy, Research Data Platform at Elsevier. And he will be taking his questions after the presentation.

**Erik Nijenhuis**: [13:30] Veel succes. [English translation: Good luck.]

**Rob**: [13:33] Dank je wel. [English translation: Thank you.]

**Rob**: [13:35] Everybody hear me well? If not, then raise your hand. It sounds a bit weird here. So indeed, thank you, Erik. I am Rob. I've been with Elsevier since 1991, based in Amsterdam. And I'm really honored to kick this off today. So I joined in 1991, as I said, and before that, I was a mathematician. And for a long time, my role in Elsevier was the Chief Content Architect.

**Rob**: [14:08] And it's really great to be back in Trivandrum. I've always loved to be in Trivandrum. It's been a long time, so it's really great to be back. Actually, I've been coming to India since 1994, and I can tell you endlessly about all my adventures in India. And I can also tell you endlessly about my grandchildren. So I have two grandchildren, one is thirteen, Emily, and one is eight. And allow me to start with a small digression.

**Rob**: [14:40] So there was a time that my granddaughter was six years old, and then she said to her teacher, "I have seven grandmothers," or seven Omas as we say in Dutch. And then the teacher said, "Impossible, you can't have seven grandmothers." And then she became angry. She said, "I set to prove it." And then she produced this.

**Rob**: [15:05] So this is a drawing that my granddaughter made when she was six, and we can learn a lot about it, from it. So first, what you can see here is that she had a really clear idea about definition or policy. So she counted the great-grandmothers as grandmothers too. So you see, that is already one important thing. Secondly, in order to prove that she has seven grandmothers, she furnished every grandmother with an encircled number. As you would say, a unique Oma ID. In other words, she created a registry of all the Omas, and she presented that to the teacher in the form of a family tree, a knowledge graph, I would say.

**Rob**: [15:55] So she didn't number the grandfathers, you see. So she really focused on the task at hand. So she didn't make a mistake that we always make in our company, to... when we go about it, we also do more than we need, then we can make errors. No, she didn't do that, only the grandmothers are numbered. And the fact that grandfathers, or men even, appear in this picture at all, is only to provide evidence that these people are grandmothers. Because in order to be a grandmother, you have to be the mother of somebody.

**Rob**: [16:31] So, and then what you don't know, but it is interesting, that the Oma number six, so in the Netherlands children usually say Oma to their grandmother, but this Oma is always addressed as Hetty. So I said, "Well, why is the box saying Oma? You never call her Oma. Why didn't you put in Hetty?" And then my granddaughter said, "Well, I didn't know how to spell Hetty. Is it with a 'y' or with 'i-e'?" And thereby, she basically invented continuous data quality or defensive programming. If you're not sure about the facts, you revert to the generic fact, something that she knows. Because if she would have misspelled Hetty, the teacher would have said, "Oh, this whole story is nonsense because you can't even spell the name of your grandmother," or, "Well, I see a misspell," and it would ruin her whole experience. So it even has an adverse effect.

**Rob**: [17:37] So by now, probably you will have believed, are believing that my granddaughter has seven grandmothers, but it is not correct. So I am here next to Oma one, that's me, and my mother is missing. So the recall, as we would say, is only seven divided by eight. So it's not perfect. But that doesn't matter, because everybody got to trust in the data. And **this trust is based on transparency, clear policies, and total focus on the business problem at hand.**

**Rob**: [18:13] So that is where we are too in Elsevier, because what we have come to realize is that **the most important asset that we have is our knowledge graph**. And a knowledge graph is not really of course Omas or grandmothers, but it is about the entities that we have. And those entities, they are the articles, so a hundred million articles from other publishers, twenty-three million full-text articles from Elsevier itself, a hundred million patents, forty million researchers, a hundred thousand organizations that are carefully curated. All these things form the nodes in our knowledge graph, and they are connected together. So all this is highly accurate, and it comes about in our data platform. So that is what I am associated with in my day-to-day work with our data platform.

**Rob**: [19:06] And the way this data platform works in a nutshell, I've drawn it here in these five steps, but you can also look at it in this picture of me when I got my PhD long ago. So the first thing is that you have to have evidence. So we treat all the things that come in as evidence. So the photograph itself is the evidence. Then the next step, so we harvest that from many places. The articles, you know what, many articles we get from many places. We get it from the author, we get it from Medline, we get it from... So we have a lot of duplication there.

**Rob**: [19:40] So what we then do is we extract entities. In this case, we recognize faces in my picture. The same thing, you recognize an entity. And in the old world we used to call this, we structure our XML files, we structure our text. It is really to extract the entities, to identify the author names, the title, God knows what. And when that is done, we match and cluster them all together. That is really the heart of our platform. And after that, so for each node we have lots and lots of properties. Not all consumers want or are allowed to get all the properties, so we create some kind of context view that is dedicated to them. So those are the steps that we have.

**Rob**: [20:21] And all this is rather probabilistic. You can also see it in the picture, for instance the head on the painting, is that also an entity or not? The algorithm could be wrong. And maybe you think it's me that on the right-hand side, but that's my father. Because this picture was in 1991, and I'm the guy on the second from the left. So this is... and generally speaking, **all these things are heavily probabilistic**. So this is what we do everyday in the data platform.

**Rob**: [20:54] And for this, **it is extremely important that we stay totally true to the source**. And if we don't, we need to apply clear provenance to the data. So I think for many years, the vendors that we worked with, they thought they do a service. If, for instance, some page number is wrong, it says 38 in the article, and it is 37 in PubMed. So maybe we need to replace it by 37. I think with that, we have already ruined the entire knowledge graph. **We need to be totally true to the source.** We don't know this lookup of the PubMed, I don't know, many errors in PubMed too, so our accuracy is on the latest percentage points. So we need to be very precise. I don't know what this lookup is. I can only accept a lookup if there is really clear measurement of the precision and recall of the lookup itself. And until then, we had better leave it the way it is in the article. And if you want to add it, then I think it should be separated with clear provenance to make a distinction that it was not the author who said it, but it was the copy editor who said it. I think now we override things, and that has, is one of the biggest problems in our platform.

**Rob**: [22:16] So what did we learn here? Slowly but surely, **we moved from what I would call markup to annotation**. And I think that is one thing that I want to stress today in my talk. So annotation means making assertions about the document at hand. And we do those assertions in RDF, in the Resource Descriptor Format. And in this way, we can observe what I call the golden rules for data. So that is what I said earlier, **we always need to add provenance.** As a mathematician I would say you have to have respect for zero or the empty set or null. When you do markup, it's really hard to mark up something that isn't there. **We need to be able to mark up things that aren't there.** So that is really important. The input is also grounded in time. Whenever we keep changing it, then we destroy it. And I think everything should be understandable through a common vocabulary that everybody knows.

**Rob**: [23:27] Most importantly, we want to allow multiple opinions about the same thing. In our model with markup, you have only one possibility to say what it is. But no, we have many opinions. Maybe the author said that it is this, maybe somebody else says it is that, and we want to have these things side by side. So that in our platform we can make the right assertion based on the weight that we give to each evidence. I think when we look at document structuring, then all these things are really very hard because you do everything in the same place. That's why we have completely moved over to stand-off annotation. The annotations are outside the file, and they are saying something about the text. So this is, **I wish we could make assertions about things in the LaTeX document, not in the XML document.** But I don't know how to address it. There is no standard language to reach in and identify... or maybe there is, and then I don't know. You can prove me wrong after the meeting, if you want to.

**Rob**: [24:32] More importantly, this whole thing about data structure that we did for many years, it is based on everything being in one file, and it totally ignores that there is a big process going on all the way from the author to publication. Many players work on the document, and they don't do that all at the same time. So what you need is a format that allows this, that many people work at it at the same time.

**Rob**: [24:56] So what I like about that is that **annotations are even more tagging than tagging**. We have been using this word tag for an XML tag with angle brackets, but I think a tag that you stick onto it like a sticky note is far more like a tag than an XML tag. So I say **we go from tagging to real tagging**, and put a sticky note onto it. And then you can have multiple sticky notes that say something about the same thing. Or you can make an assertion about things that overlap like this, which is not possible in XML, that you have things that are sort of nested inside. All these things are easily possible if you do it on the outside, and we can add even more assertions over time. So that is important.

**Rob**: [25:47] So yeah, it's been a long journey. I'm privileged that I have been with this whole thing in Elsevier for 30 years. So I joined that particular team in 1995, and I should pay tribute to all the people who already did that way before in the 1980s, who decided that we should go the way of SGML. I mean, there was a lot of debate at the time if that is really the right thing to do. And I have here some dates, I won't go through the entirety of the dates, all these milestones here.

**Rob**: [26:21] So one thing we did in 2015 is that, you can't see it here because the top dot dot dot is there, but anyway, in 2015, we had worked on what we called an enhanced reader. So on our main platform, ScienceDirect, one thing that we tried to do is to improve the way you read an article. And we experimented with many things. At one point we tried to embed the PDF on the page on ScienceDirect. And we hacked that PDF so that it became active and that you could click on authors that weren't tagged before. So that never really worked well. And we also tried the other way around, to make the XML look much like the PDF and then look at that. Also we didn't get there.

**Rob**: [27:18] **My wish would be, I'm going to say that 20 times I think today, is if you create a LaTeX file, that you would create an HTML file and a PDF file in one go and they are more or less the same. We can achieve that, then we are there.** Now I thought we were not there at all, but then I read in the program book, and I read this whole story from Radhakrishnan, and I think, "Oh my God, maybe I'm telling something that is actually already long there." But now I think it's still not precisely there the way I want it. So that's good.

**Rob**: [27:51] So I thought, we are now in 2023, and I have these 30 years of experience in the articles, around the content architecture, as it is called. So let me share with you then a number of reflections of what I concluded, what we did, and maybe what we could have done better in all those years. I call them the reflections. And I have to watch the time, but yes, I have a lot of reflections, actually this is pretty good, I can go on for hours on these reflections.

**Rob**: [28:26] So one thing, so I think, well, the first one is, what did we say in the 1990s? People had really profound ideas about what we wanted to do with the data. And I think still some of them are still very important. So if you read that article from Radhakrishnan, I think the draft is in the program book, you see how important it is to preserve your digital archive. So I think to preserve your digital archive is really good. You don't want to have the content and presentation so closely tied...

---

**Rob**: [00:00] ... it looked nice, but it was ... I don't know if it was needed. And I'm here to say this a little bit provocatively, I think.

**Rob**: [00:13] So another example is, yeah, so we, I think the DTD that we created was heavily influenced by people like myself who started in a role of a copy editor, desk editor as we call them. So at the time it was common that a desk editor would indicate where in the article a figure, a float, should appear. So then we put in, at the time that it was still paper, we would put in figure three with a red circle around it, and then the typesetter would know this is whereabouts. So in our DTD, we have CE float anchor. That is the place where the float should come in, approximately. So yeah, that somehow allows us to give this indication.

**Rob**: [00:59] But we don't have anything like this for table column width, for instance. Table column width used to be the total remit of the typesetter at those times. So we didn't even think about saying how the table columns should behave. But now we see that's really a problem. And people in Elsevier have been told, "Oh, now we have LLMs, maybe we can automate it." Oh, maybe for 20 years, 30 years, people have been trying to automate it. But what they should have seen is that **the DTD was wrong all along**. We should be able to indicate this. And this is where we should ... So we did do it for the float anchor, but we didn't do it for the table column or the math breaking. A really, really hard problem. And everybody else realizes that, except us. For instance, if I do a lot in Confluence. I don't know if you know Confluence, I do all our documentation. Of course, it provides an option to adjust the table columns and it will stay the way I like it. Those things should be possible but they weren't possible. So I think these are kind of things where we made a mistake.

**Rob**: [02:08] People were even so pedantic about it, I would say, that I mean there were ... I don't want to say anything bad about my predecessors, but in the 1980s we had a predecessor of our DTD. And then there wasn't bold and italic in the DTD, but we had emphasis one, emphasis two until emphasis nine. Because those people said, "Look, emphasis is ... " and they had this vision that you could also show the article when you swap around. It was just about the emphasis, not that it is bold or italic. That was just a style element. But I think nobody, no author, nobody would really accept anyway that suddenly the bold and italic would swap around in their article. It's just a silly thought I think that was there. And they said, "Yeah, no, but when it is italic, they should really, really indicate, for instance, *E. coli* that is a genus name for a species. So we shouldn't say italic and italic around *E. coli*, we should say species and species." That's what we want to do. Because that never worked because how could we have a system in which everybody ... if the author tells us that's fine, but I think **to add an interpretation can only go wrong**. So I think it's far better if we just use italic and we annotate on the side if we want to, because we didn't even use it, and say what it is. So that is how ...

**Rob**: [03:41] We also said number two. Yeah, we are going to republish the content in countless ways. So an article will also appear in a book and it will appear there and it will appear there. So that also really never happened. And if it happened it was not exactly the same article. It has its own DOI, its own citation details. So then you can take a copy, but it's not exactly the same file. And if you look at what the files look like today, they don't have all that metadata, because if you would put the data in, it wouldn't be republishable. So I think that makes these files completely useless. So I think that is another thing that was a good idea at the time, but we never used it.

**Rob**: [04:26] And then we were extremely dogmatic about XML first. So that means that if you have an XML file, from it you get both the online rendering and the PDF. And yeah, I already mentioned a few things that make that hard, and it also relied on extremely heavy standardization. Now I think in the end maybe it was good that we did that, otherwise we would never have been making any progress. But now all that is out of the way, you can wonder, is it really that important? I was yesterday reminded by Kaveh [?] that we introduced LaTeX first. There we already deviated from this. I looked it up when it was, it was 2007, so that was two years after the last big generation DTD 5 was implemented. And I think that time, 2004, many people here in the audience are as old as I am and worked on those things too. I think in 2004, 5, the things that we did then, we went from the sort of starting point of our electronic world to the well, actually what is still there. That was a fantastic time to think back on all the things we achieved. But we did all those things and now in hindsight you say, well, we didn't really exploit that so much.

**Rob**: [05:48] One thing that has become extremely hard, I come back to that on another slide, is that if we even want to have the tiniest change to our DTDs or schemas, then it's extremely complicated. So that is ... And then another thing that I can repeat is that for us, **it's really no longer about the article, it is about the entire corpus**. If you want to have this knowledge graph working, my attention is to 100 million articles, not to this one single article that we're looking at. And therefore you want features that are ... no consumer of the data wants a new feature only trickling in in new articles. If it's also present in older articles, people want to see it there as well. So the thing we need to be able to do with all our standards is to support that. Not by going back to all those files, redo them, but to add things to it after the fact. So those things were really sort of in another line. So all these things were in some sense true that we wanted it, but they never turned out to be so relevant.

**Rob**: [07:04] So another thing that I thought, I mean you all probably know the expression, you should eat your own dog food. Now I think what is good about the LaTeX community, and that's a good thing, eating your own dog food, I don't think that it is bad. So I think **in the LaTeX community, eating one's dog food is one of the best**. And then there is this whole community, there is the CTAN, I mean I think that is all unrivaled in the world. And I think other people could ... I think it's sad that not the whole world knows about it because I think that's the best that is. So but did I always use my own dog food? I was Chief Content Architect of Elsevier and I was responsible for the Elsevier DTD. I also created documentation together with my colleague. So these were two books, tag by tag, first for HTML and then for XML. The green one was 444 pages. Now, to be honest, I wouldn't even know how to do it in my own XML. Well all I could do maybe is do it and then ask one of the vendors who are many of are present here in this room would do it for me. But I want to do these things myself obviously. I have this book and so I think there is no tooling that is widely available. It's really hard to write these things. Oh right, should I write all my angle brackets in Notepad++? I don't know, I have no tooling. So of course we did that in LaTeX. The whole book project was done in LaTeX and there is, I don't know that there is any other way.

**Rob**: [08:44] And to have another point about this. Here at the bottom there is a snippet of a newsletter that I've been publishing in Elsevier, where many people subscribe to it. It has been around since 1996. And then there you see in blue, that is a third-order heading. Now yeah again I would say I'm Chief Content Architect of Elsevier, I'm telling everybody you have to be very, very precise on the tagging. But I have never in all those 30 years, 29 years, coded that in Word as a third-order heading. I just made it italic and blue, over and over again. And this worked. So maybe this is sort of a almost ... yeah so you could see this is a big disconnect. And actually I've never got it to work in Word to have a heading that has run-on text instead of a new paragraph. It works but it's very hard, I think. So you say, "Well no, of course it's better to have a third-order heading." But it has never, in any way, interfered with our ability to read those files or to deal with them. So I think that is somewhat overrated. So we did this, this has worked for all those years, and everybody was happy.

**Rob**: [10:03] So what would I do if I would write documentation again? Actually I wouldn't know. So I have been trying to do that with well, maybe it's fantastic to do it all in Markdown. I like Markdown a lot. But then you, it's nice, it looks like it can be quick. But actually is it so nice? I don't know. I can't have a bulleted list in a table cell. Well you can if you then before you know it you're starting to write HTML files in the way we used to create our own homepages in the 90s. I mean that can't be the answer to things either. So there is no macro language whatsoever there. So if I want to make sure that a tag or a predicate, as we would say, let's say the title field, if that is clickable, comes in the index, then I want to have a command that deals with that for me. But there is nothing available. I can invent it myself, but that surely shouldn't be the way forward that I'm going to invent my own language to do something that everybody needs. And then there is something called MkDocs, it's quite beautiful, it gives you this whole website. But then every user of it needs to install a server. People are not allowed to do that if they're not an administrator. So I think that's simply not working. So that's why I think, so you could say, is then not obviously the conclusion that you want to do it in LaTeX again? I'd say well no, because what I don't get there is right away a fantastic like say Confluence space like construct where all the things are also in a website. So if I could with the same LaTeX file do that, I would be in business right away. So I think that is I think generally speaking the thing that I would ...

**Rob**: [11:50] Here's another note. In fact, you could then say, yeah, should I eat my own dog food? No, I've never eaten that. It's a bit shameful maybe, but at the same time it was not really that relevant. Good. Another reflection. So another thing was that we, yeah, it's nice that in an XML file you have DTDs, schemas. So what you can do is you can validate. And I think that's very important. I think that's also something that I think is missing in LaTeX, that you, I don't know really good ways to have a ... A markup language should always come with a way to make sure that what you want to have is valid. So but what we did was ... So I have a few comments to make about that. So first of all, we can be so proud about it. I also said, "Well, you get a nice DTD available." But when was a DTD ever sufficient? So I'll tell you something about one tag in our DTD. It is the `ce:cross-refs` element with plural, okay. So this is a one-to-many link. So for instance, if you have figures one to five, then you can make a link from the text "figures one to five" to all those five figures. So you have five destinations. And the way this works, I was about to say nobody implemented it, but I know that there is one person here in the room who did implement it. It's that there would be a drop-down list of all the five destinations. That was the mental model that you should have. And this drop-down list would be constructed from the labels of those targets. So if you have a figure, it's called figure one, and you put it in the drop-down list, then two, three, so you pull it from the targets of those five cross-references. So in order to make that possible, there has to be a validation that whenever there is a `cross-refs` plural with those five destinations, that all these targets have labels. So you can't express that in the DTD. Not possible. So therefore we always had a V2 that the person who is managing that for, now for many, many years is also here in the room. We need that on top of the DTD. So you can say DTD is great, but actually they've been grossly insufficient.

**Rob**: [14:31] And then what happened was people didn't follow the rules. But of course we never told them what the rules were, so you can't blame them. And the rules were really only the producer of data should do really, really good validation. The recipient shouldn't. Because I say how can he say that, I mean they need to be sure what is right. Yeah, but if they do that against exactly the same schema, that means and that is why we are completely stuck right now, that with the tiniest change, everybody also needs to install that new schema. What they should have done, create their own data model, pick from the file what they need, they can check whatever they like about that, but ignore everything that they don't know. So I think that is an extremely important principle that we should have in the future. We wanted it 20 years ago, but this is not how it came about. Everybody creates their own model of the file based on exactly the same schema and then innovation is almost impossible. So what you want to do is to do a very clear validation on the start. So the producer needs to do all his own... these were his own health checks. So that is symbolized with this orange thing. I call them the blood tests. You look... if you go to the doctor, there isn't... you get a often a blood test and it has... it shows you well certain values are above or within ranges that you want to do. So the ranges are known and the values are known. If it's outside the range, it's still not sure that you're very ill, and if it's in the range, you're still not sure that you're healthy. But at least it gives you an idea. So all these things need to happen before the thing is sent and in the end, you put the really critical things, that is my red cross, needs to be done there, doesn't even... shouldn't even leave the producer. The recipient however, they what they need to do is only focus on the things they need. And the trouble is what today every receiver does, is repeat what the producer has done. So if we can somehow change that, and I think that's now no longer possible with the things that we have. And also I think the whole corpus that is robust enough to allow a few errors as long as you fix them very clearly. So this whole thing about absolute... the things being absolutely right is a bit... yeah you could say that again it was a very true thing, but it became, yeah, it made us more and more irrelevant. We can change that, that would be good.

**Rob**: [17:19] I'm going a bit from topic to topic, just to show the kind of things that I had in mind. So another reflection is, so I have been in the world of data modeling a long time and many people came and went that helped me with that. And I think I see two approaches to that. Especially if you come from university, you have a lesson about data modeling, then my experience is people are not modeling the data, but they are modeling the words or the world. And I think the last thing you should do is to be very pedantic about knowing what an affiliation is. So then you have people tell me an affiliation is a relationship between an author and an organization. And yeah, that is true. Unfortunately, how people use the data is not like that. If you look at, across 100 million articles, you see some of these examples. So I, when I came, before I came I thought maybe it would be fun to show you all sorts of things in live demo, but I didn't know the way this conference is doesn't allow that very easily. Here are some examples on the screen. So for instance what you see here, the top. That's what people put in. Suppose you have backslash affiliation, some affiliation with angle brackets and then it says "Associate Professor of Pathology". So that isn't an affiliation, that is a job title. Yeah, that is terrible.

**Rob**: [18:57] And then this second example, this gets even more ugly. Then somebody gave as their affiliation the "Departments" (plural) "of Medicine and Cardiology in India". Now, that doesn't... doesn't tell me anything. I mean, in India, come on. How is it possible? Nevertheless it appears like this in an article. But so I don't want to fuss too much about the fact that it just says India, I wanted to say these are two departments, Medicine and Cardiology. It says so, departments plural. And then the third affiliation is not even an affiliation, it's a street address. Okay? So that is, now yeah, so if you are a data modeler, you're deeply shocked about that. At least if you do, if you model the words. But **I don't model the words, I model the data**. And this is what people do, so let them. And it needs to appear on the place where the affiliation is. Fine. We also don't need to break it up in five pieces. That's what Elsevier people have told suppliers, "No, you have to make it not A B C, but A B C D, because there are so many departments in there." Not, not important. True that these are two departments, but irrelevant that we have to split it. And that is all thanks to the fact that we don't structure, but that we annotate. You can annotate this "A" with two organizations, no problem with that. You wouldn't be able to do that. Same with people. So you have, for instance, it can be the, so what people call an author is in general really just the byline, who it was written by. For instance, you could have Pierre and Marie Curie. Okay, that is two people. But it doesn't say Pierre Curie and Marie, it is Pierre and... so it's a contraction. Again let's put it in. I don't care that that is not truly one author. It's a byline, we can always annotate it. And I think, so many of the constructs that were created in our data model have been based on the idea that what is in an affiliation or an author is an affiliation or an author. But in fact, it's sometimes zero or many. I have zero and I have many. And here at the bottom is even a reference. Now, there is not one reference, it is a whole sentence. The sentence with, well, how many references are there in this one sentence? One, two, three among others already starts. One book, another book. In total there are three here at the top. The bottom one is one. And this is a style that people use. Who am I to say that that style is wrong? So that is, that's perfectly okay. So this is not the author's problem that in my DTD I have made it three or four different things. I think it's really hard to, to deal with this because somebody thought that the reference is always one-to-one. I think that's another proof why we want to annotate. You just style it like this, everybody likes it, that is what the style is. Nothing wrong with it to my mind. And you can indicate how many references you find there and what those references are. So that is another sort of motto that I have. **Modeling the data is modeling the data. That is it.**

**Rob**: [22:24] And then one more thing. I'm reaching the end now of the, almost of the reflection. So another thing that we keep forgetting is that an article is not alone. And with that I mean it, yeah. If we look at one XML file or one PDF file, one LaTeX file, or maybe it's a whole set of LaTeX files that produces this article, it's all from the author's perspective. And maybe from the typesetter's perspective. But in our process, a lot of things happen completely independently. For instance, something like a copyright notice. It completely follows its own workflow. So that is just an interaction via a website. So before, you didn't know it, and at some point you suddenly know it. Now, for years the whole workflow was stuck then. Should we then because the source of truth of the files was at the beginning, people have to add it in and then push it through. These things should just be changeable in their own right. So we came up with an idea to allow stamping in the PDF. And this is, I think a bit of a kludge, but it works, but it should be really sort of officially supported, that you can say, well, the copyright notice comes from another direction and we can just put it in. It doesn't have to come from one place, but from two places. Should be possible.

**Rob**: [23:52] And another thing that we are not really well dealing with is the fact that multiple revisions come in. That happens a lot. So it is terrible. If you have an article, the author and we process it, now a lot of work goes on in this whole process. That is fantastic. But now a revision comes in. Of course if the revision is huge, fine, but if the revision is tiny, then still it needs to go through all these steps. And maybe we even pay the same price for it. I think that is not right. I think there should be ways where you can really have **everything is optimized for differences rather than for new articles**. I think one way that I like very much to do that is to make use of... one of my favorite things, I'm unhappy that I didn't take the time to put a patent about that, is the reproducible identifiers. They, they can be helpful. A reproducible identifier is a hash of the content. It almost makes no sense. You take content, you have to be a little bit smart to do that because it needs to be unique. And then this paragraph ID you create it from the text of the paragraph, maybe a little bit of its place as well to make it unique. And what good is that? Because the tiniest comma, and it would change. Yes, that's the whole purpose. The purpose is not to get everything right or to get a unique ID that works across all versions. It is to detect which ones are different. And all my annotations that I made about this paragraph will remain valid if a new version comes in and the paragraph had remained unchanged. I can just keep on referring to it, which normally there is this sort of, how do we call it, master-slave kind of a problem. That if you have an article in the way I envision it, an article with thousands if not hundreds of thousands of annotations on the side, at the moment a new version comes in, all these hundred thousand annotations suddenly are invalidated because you don't know for sure anymore what the text looks like. So they work extremely well. So I was, when I said... I keep always saying something about this fourth paragraph, is about this and this topic, or this fourth paragraph contains information about this chemical compound. These are the kind of annotations that are there in about this paragraph. And the moment you get another text, now yeah, if the comma changes, maybe that has a significant... then you do it again. That is the price you pay. That is what we are there for, to recompute, but we shouldn't recompute anything that wasn't legally necessary. And then, I said it already earlier, the entire database matters. So I call that the ocean and the waves. The inflow are the waves. To navigate the ocean on the waves is already very hard. There's storms, there is high waves, ships sink. I mean that's the process that many people are on. But underneath is four kilometers of water mass. That is my corpus. All the millions of articles that are out there. And they need to be correct too. They need to be constantly refreshed because they can't go stale. And another thing is a new data point can only be good if it is useful across the corpus. If you have a data point and we are so happy that we have it in this one article, it makes no sense to me if it is not consistently done. And there I want to go back to the picture of my granddaughter. She had it right in her mind, she would never edit a field that wasn't consistent in all the corpus. **Precision, precision, and recall are everything.**

**Rob**: [27:30] And then lastly, this is my last reflection. Big debate in Elsevier, maybe everywhere I don't know, is should we abolish two-column journals? I wish we had a whole session about it. Should we... if only we could abolish two-column journals. So now let me say here something shocking: **Let's embrace two-column journals.** That is just to make people slightly unhappy. So what I did here to illustrate that is I took a page from ScienceDirect and I expanded it to my screen at home. Okay. So then I get this. So this is absurd in some sense. It is almost all white. But is it absurd? No, it isn't. This is what everybody wants. So you have this column. We should decide on a width. If you look everybody else does it like... too... if you see how people implement blogs or I don't know, then the width is... and you can scale it. Not that you widen it and then everything starts to reflow, no you settle on a width. And everybody settles. If you do that, then maybe we can even sort of get far closer between an HTML version and a PDF version. Nobody really wants it to be so incredibly different. And after 20 years of experimenting, I think the product people from ScienceDirect, who interacted with thousands and thousands of consumers, realized that that was the case. That if you widen it, it shouldn't suddenly go all the way like that. Nobody wants it. There is a sort of natural width in it. And that would also solve all our problems with math breaking. Because if you resize, suddenly the breaks will change. Going not going to be really, really nice. Oppositely too, if you have a very, very narrow column, and you have do a lot of work to split an equation array in smaller parts. Now if the text is right, then it just looks silly that you have text and suddenly this tiny little formula. So I think it would also, if we would just forget about all this reflow, and say, well we just have one column. Now if it is an extremely long column, and then in a two-column, you can just put them next to each other. I mean, I'm saying it deliberately simplistically, but that is just to make the point. So I think these were my reflections, so I'm nearing the end...

---

**Rob**: [00:03] I can sum this up. So if I were asked to create a new DTD, a next-generation content standard that would last us for the next 20 years, what would be my requirements? Now, I don't know if this is even an exhaustive list, but I thought I'd sum up what I just said. **First of all, it must center around annotation, not about markup in inline.** We just have some sort of HTML pages with as little information as possible. **Then it separates content from annotations, so not really presentation from style, but content from annotations.** It stays really close to the display. This whole idea that you have countless ways of presenting it hasn't happened. Nobody wanted it. And all these annotations are extremely strict. By that I mean you can validate them precisely. But as I had in my example about the footnote references or these affiliations, the way it looks doesn't matter so much to me. That is what I call **flexible content with strict annotations.** And it always must come with a language that you can allow to validate that what you need is there. Because our whole mechanism cannot work when everything is sort of half-hearted. We need to know, if we expect this particular... my example of the label that needs to be there for a Crossref, then we need to be able to be absolutely sure that it is there. Otherwise, the product doesn't work. It has to have this idea about many people working on it, so call it stamping, call it what you like. And then it needs to be able to incorporate content inferences.

**Rob**: [01:55] Now we have such a standard. I must admit I'm not entirely sure yet if we should go there, but it is there, an official NISO standard, and it's called Content Profiles Linked Data, CPLD. You can look it up on the NISO... [inaudible]? NISO. And it has two components. One is canvas, as they call it, that is the HTML5 rendering. Of course you can do an HTML5 and then you can have extremely nice typesetting. Yesterday at the drinks, everyone was going on about perfect typesetting. Don't get me wrong, I love perfect typesetting too. I even love the funny... since I am friends on Facebook with Radha Krishnan, I follow the funny posts of these cartoons. And then still it touches me that it says 18 to 21 July with a hyphen and not with an en dash. Oh, but I go over it. So I like perfect typesetting too. Don't get me wrong. **But nevertheless, we need to be able to render a whole website in HTML5 also. And then all our annotations in RDF, and in particular we do that in JSON-LD.** That is what this standard is.

**Rob**: [03:12] And lastly, it fits almost all these boxes I think. You may say, can you really go without the parsing of an XML file? And I say, no, yeah, no, my annotations are not just about semantic annotations. I also really meant the structure of the article. Here is the section, here ends the section, all these things we put outside of the file. And then we can use SHACL schemas to check them. Because we do that... so **we don't do that with DTDs anymore, we do that with SHACL schemas in the JSON-LD.** We let the HTML do what it is. Don't get excited about putting even more nested divs into the HTML file. That's what many people do. Div, div, div... and you might as well create a new DTD because that is just silly that you use divs and you can't validate them with a class. Then you could better have created a DTD where you can validate it with a schema. No, leave it alone and do your validation in a SHACL schema. And **in a SHACL schema, you can make SPARQL queries within the graph. Now that is so beautiful.** You can do a SPARQL query, so my test about the Crossref can be done there. You just check, you do a query, say do all the destinations really have a label? Yes. Then we can push it onto Crossref.

**Rob**: [04:30] So the difference is we don't... well, we maybe still need a [QA tool?], validation tool, no, but it's far better we can use an official language for that. SHACL schema. So I think that's always better if you use an international standard rather than some homegrown system. So all these things tick the box. But there are also many things that aren't so clear. For instance, will it be extremely costly? That is what people said when we implemented SGML. Is that... I mean everybody has so much tooling for what they have. And steeply embedded in our pipeline, so it's really not easy to... we would all need to refresh our documentation. But I think that is the only alternative that I currently have in mind, and then what I'm hoping that in the coming days also after the... we're here for three days, maybe if people have completely violently opposite ideas about this, I'm ready to hear them out. And maybe some of the things I said that weren't there are actually already there. And then I'm also really keen to hear it. I'm also really, really keen to hear all the other... the program looks super fascinating. So I thank you for your attention, and 30 years of reflections about content processing in Elsevier.

**Host**: [06:05] So, does anyone have a question for Rob?

**Audience Member 0**: [06:12] Thank you, Rob. It's a great presentation on the 30 years of reflection. I'm just thinking CPLD, in the AI world, how is CPLD going to be beneficial? That of... now the world is changing, not ready to read the whole article or whole book or something like that. They want to go into the particular topic or search that... example, ChatGPT raising the question. How we are going to connect CPLD with this particular thing?

**Rob**: [06:39] Yeah. That is a good question. I think there are two sides to it. So one side is that with LLMs and so on, really you see how easy it is to structure data. So you can give it a format and it will already do a lot for you. What we also see is that actual text has been increasingly important. So one product that we have is Scopus.com, and it contains all the abstracts of articles from 7,000 publishers. And those abstracts were actually not really so looked at so much. We were so concentrated on creating our knowledge graph, linking articles to authors and so on, that we neglected the abstracts a bit. Now this has completely changed because we query these LLMs and then give the context from our own articles, and we call that the retrieval augmented generation, or RAG. You've probably heard about that. And then we let the LLM format the answer out of that context.

**Rob**: [07:46] Now the tricky thing with our articles is that it's all scientific stuff, math formulas and extremely difficult concepts. And if you give that to an LLM to read over it, they often don't get the math and they don't get all the scientific concepts properly, and they garble it up. And so the LLM basically hallucinates an answer. And **to prevent that, the text must be extremely high quality.** Actually what is now our biggest worry is that all the things that we want to do is present the information inside those abstracts. And so I think that is our answer to AI. So I think our own AI offerings need to be able to look at the data, ideally therefore to the original data, and in my idea that would even be far better if we could look at the LaTeX file rather than some extract from it. So that should be the core, maybe if it's a LaTeX document. So I think no, the... in the AI world, the text quality needs to be very high.

**Rob**: [08:00] And then another thing that's there is that yeah, I think, or what is generally said in the management of our company, is there's a lot of opportunity for other people to be same players in the world that we have been, thanks to AI. But I think the difference that we can make is that we see that all these LLMs and what not, they all get... that **we can make a difference thanks to the large knowledge graphs that we have.** And to work on that, to get it super precise. You can get open access content really easily. But the only way we can do is to make sure... make it clear that there are so many issues in the data that we don't have, thanks to 20 years of the kind of processing that helped for us. I don't see that go away so easily. So maybe one more question?

**Audience Member 0**: [09:00] Hi Rob. I was wondering if you would have anything to say about publishing in multiple languages, any challenges regarding annotation, tagging?

**Rob**: [09:16] Yeah. So any... thing in particular? So I think of course, I didn't say that on the list, but all our standards, first of all I'm starting from the starting point that Unicode is completely embedded everywhere. So different languages, different scripts shouldn't be a problem anywhere. So that is point one. So I think all the things that we should do should... yeah, maybe that should be a bulleted list for a requirement for content standards, is that we can... we have no problem having multiple languages. But if you think about sort of articles that are themselves in two languages, or I'm not exactly sure what you have in mind to... to support languages I think should be a given, that that is possible. Actually, I think theoretically it's not a problem right now even. So I think that would be... but if you think about more fancy things like articles that are in two languages...

**Audience Member 0**: [10:20] I'll give you an example. So in English you would say "page 3 out of 10." Some languages' syntax would go the other way. 10 would appear first and then 3. Syntax could differ across languages.

**Rob**: [10:36] I see. No, yeah. I understand now what you mean. Yeah, so I think now that was... I should have listed that as the fallacies of sort of generating text from the tags. So I think that is true. So **it is almost impossible to deal with all the combinations and permutations that would be there.** So no, I think that is why I think my... what I thought should be my new idea, and my new vision, is that we don't worry about it. We have the text in its textual form. That is my point. So don't even try to capture it. It just says it in the other way around or however it is in that language. But now, you point at it. That is the whole thing about annotation. You point at it and that assertion is in our standard vocabulary. It says it is page 3 or something like that. So there that is not in a language, that is in its own language as it were, it is in the annotation language. So I think that to... that is why I say **separate content from annotations.** That is precisely the thing. Yeah. So an extremely good point. So I think yeah, you have these things. So I hope that gives a bit of an answer. Maybe one more? Oh, [SKV?], that's dangerous.

**Audience Member 0**: [11:56] Very nice that you have something created 30 years back... 20 years back... still come around to talk about it and still relevant. Very rare class of... One of the problems the SGML DTD has been that it brought that HTML elements still in it. Like you have the DTD and the doctype declaration. So it still got its roots in the old world, not moving up to the new world in the sense namespaces and modularity. And also now that HTML5 has come, W3C standards like MathML. Because W3C MathML is slightly different from the W3C standard. So it's not nice to have two DTDs for the same standard, right? So there are still elements of its sort still managed. Because if your MathML implementer... MathML is different from W3C standard elements... So I'm saying, how much of Elsevier's XML is still integrating with the standards in the sense of the... W3C standards also have moved from XML to semantic web and all that. So there is also movement, right? Movement and agility, right?

**Rob**: [13:42] Yeah, so I think the... I'll give a short answer because the time is almost there. Yeah, so I thought... so I think one way that I interpret your question is that instead of going in a completely different direction, we could just sort of modernize and do all the things that we have left behind. So I think the standards haven't stood still. And so that would be an option. But I think that to make adjustments to the present standard... the present standard is associated with so many conventions that have been ingrained in people's head, that **it is better that if you want to make a step change, to really make a step change, rather than to try and fiddle around with what is there.** Because people won't... I mean it's almost psychological. The moment you make this sort of other change, then suddenly also things that have nothing to do with the schema will be open in people's head to change them. Whereas if you stick with the old schema, and you make some adjustments or modernizations, then I think they will apply... everybody will apply all the conventions that have been all there. So I think there is both a modernization aspect, but also a pedagogical aspect in that when you change, people suddenly feel more free to think about new ideas. So I think that would also be an answer. But Eric, I think it's... my time is up.

**Host**: [15:13] Yes. Thank you, Rob. Give an applause again to Rob. So our next speaker is Eric Brown. He's going to talk about the Comprehensive TeX Archive Network. Good luck.

**Eric**: [15:48] Good morning. This talk is about the current state of CTAN. At first I want to make a little survey. Who of you knows what CTAN is? Please raise your hand. Oh, great. And who of you visited the website from CTAN in the last four weeks? Aha. And who of you installs his or her packages from CTAN? Oh, I'm surprised. Okay. **CTAN stands for the Comprehensive TeX Archive Network**, and I'm proud to say it's almost comprehensive and it is an archive now. The archive is not publicly available yet, but it will. And TeX and network is self-explaining.

**Eric**: [16:50] The most people know us if they visit the webpage, ctan.org. And this webpage contains a catalog which contains almost, I think, everything about TeX, and it contains a repository and files, which is slightly less than the catalog. I come back to this later. And what CTAN does is the upload management. That means creators can upload their software. We have the catalog, this is what you see via the webpage, and we send out announcements, and we discuss with authors and so on. And additionally we are the place where more than a hundred mirrors all over the world synchronize the software packages for faster access.

**Eric**: [17:59] Most of the packages are free software. These are the packages which are included in TeX Live. I don't know much about MiKTeX, so I only talk about TeX Live, but MiKTeX may be have the same use. Okay, and the last point is **the packages are the sources of the distributions, and the distributions are the very actual... at least this is the case for TeX Live, so I'm wondering that you are installing from CTAN directly.** But for most packages you just have to wait a few days and they appear in TeX Live.

**Eric**: [18:48] Okay. This image shows or tries to show the connections from CTAN to the outside of the world. I recommend to download the presentation from the conference's website. Yeah, it's a little hard to explain everything now. The grey part is CTAN, and the outer parts where the little icon with the people are the users, the package authors, and the purple part are the distributions. You may notice that there is a loop between CTAN and the distributions. That's the case because **CTAN hosts also the distributions, but the distributions take the packages from CTAN.**

**Eric**: [19:43] Okay. The history of CTAN goes back to at least 1993. That's the earliest reference I found. An article from George D., George Dennis Greenwade. Don't get me wrong, the article isn't boring, or CTAN isn't boring at all, it's just not the topic of this talk. I recommend to read this article because what you find there is to large parts this what is CTAN today. The main difference is that the original concept from CTAN was made for several main servers, and since long times, at least since 20 years, we just have one main server in Germany. More about the eventful history you find in the English Wikipedia page.

**Eric**: [20:44] And now about the current state of CTAN. The core team of CTAN consists now of six people. The upload managers, the people who install the uploads and discuss with the authors. This is Petra. She's a retired teacher. Manfred is now a retired software engineer, and both do CTAN since long times, I don't know how long. Then it's me. I work as a systems administrator for Linux at the Jena University in Germany. And we have Vincent Goulet. I was able to recruit him at the last conference in Prague. And he's quite active now. He's a professor for statistics from Quebec in Canada. And he's so active, if you read the announcements, you will already have found his name.

**Eric**: [21:50] The webmaster and the programmer of the... how he calls it, the portal, is Gerd Neugebauer. He's also a retired software engineer. And recently Oliver has joined. You may know Oliver Kopp as the author from the famous JabRef program. The system administration is done by Manfred and me, and the mirror management is mainly done by Petra. Mirror management means some mirrors retire, some mirrors are added, and is just a few entries in a database, but everything has to be checked. So the HTTPS certificates and so on.

**Eric**: [22:48] Okay. We use hardware of course. As I said, for a long time we had just one server in Germany, but it got unmaintainable. So we had to split it, and we took the chance and split it to smaller servers which are specialized for its tasks. The main server has a lot of storage space and memory. There runs the portal, which needs really a lot of memory, 10 gigabytes. And we have the mirror server, or the mirror redirection server, and this server has to be fast. And so it has a fast network connection and fast storage. And the original German machine had some more servers. In the parentheses, name services, webpages or services from the German TeX user group, LaTeX Project, and so on, but that's not the topic of this talk. It's moved away. We now have two servers for CTAN.

**Eric**: [24:08] Okay. This... apart from the MaxMind GeoIP geolocation software, which maps IP addresses to locations, we use just standard free software. But we had to program or write many applications and scripts by ourselves. The portal is written by Gerd Neugebauer in a Java dialect. I don't know its name. And the mirroring is an Apache handler. This is a part of... included directly in the web server Apache for fast access. And for mailing lists and the mailing list archive we also use standard software. This is Postfix and Mailman in our case. And you see this movie-trailer-like screen. Go on the Uniform Resource Locator credits and you will get every detail you may want to know.

**Eric**: [25:28] Okay. Where are the redirects sent to? We have more than a hundred mirrors all over the world with hotspots, as you can see, in North America, China, and Germany. And there are also some grey countries. The light grey means that there is no mirror at all. So **we have in South America and Africa all together only five mirrors. That means if you live in these countries or you have connections to there, try to convince them to open a mirror.** It's not that hard, I come back to this later. And I'm happy to say that this map from the last week is already outdated. Because on Monday this week we got a new mirror from a nice guy in Sofia and is our first mirror in Bulgaria.

**Eric**: [26:42] So, how you become a mirror is described on the CTAN website. Basically you need a permanent internet connection, and a fast network interface would be good. And you need a piece of software which synchronizes the data from the main server. In Linux it is usually rsync. I think this program is also available for Windows, but I do not know. **It's crucial that mirrors synchronize at least once per day.** This is an old rule. I think you can do it every six hours or more often. But it makes no sense to do it more often than once per hour because the portal is renewed every hour once, so... once per hour is good. The mirror redirecting server is fast enough. If you run a mirror, do it so.

**Eric**: [28:00] We don't check the correctness and completeness of all files directly. This has to be done, I know. We just check a timestamp in the top-level directory, and from this timestamp we derive the actuality of the respective mirror. And this actuality is displayed on the MIRMON page. If you go to this page you see some are indicated as green. These are used for the redirector. Some are red, these are not used, but we hope that they will sync the next time, they are still available but not used for this time, they are indicated as red.

**Eric**: [29:13] So as I said, for the uploads we check only for formal correctness. Formal correctness means for us file format, directory structure, license is very, very important, version indicators, which has to increase with every upload. That means we don't judge for the contents, except in some cases where someone loads his personal abbreviations. And of course we use programs for checking, it would be too boring to do this every time by ourselves...

---

**Erik**: [00:00] Mainly we use the program package check. It's available on CTAN. **If you are an author, please use this program if package check says your package is okay, we will not find any error either.** And my experience is that users of L3build from the LaTeX3 project also send very good packages from a formal view. So I don't know if the users of L3build are used to make good packages, good in the sense of CTAN, or if it is L3build, but use L3build and your packages will be good, maybe.

**Erik**: [00:58] **CTAN is always evolving.** Since a few years, Gerd is programming the new webpage. He calls it code name 3.0. Maybe now it has another name, the name is a few years old. And he replaces outdated technology with modern frameworks. And especially the user face is cleaned and is made for smaller and bigger screens. And he will integrate the historical archive. It's not really historical, it reaches back 15 years, I think. He will release it in this year, I hope. And details about this, Gerd is very open with his work, you find on the build page. There you find links to the source code at GitLab. **The source code is open. You can join it. You can try the new portal software if you like.** And Gerd wrote very extensive documents. The UI design guide is about 300 pages and there are many pictures of the upcoming version of the website.

**Erik**: [02:35] What CTAN does is a recurring routine. It's not only the upload management, but **we have to find for new packages a concise description. This is the little text you find on the CTAN page of a package, and a one-liner which teases, and suitable topics.** And if you are a package creator, we would like if you give us these three things: description and topics, because we are not firm with every type of packages. And sometimes it's hard for us, if it's computer science or specialized Chinese typography or so, to find the correct topics. And we have to check the links if they are still working, or some packages refer to other packages, if the other packages still exist or are still in use. And another time I emphasize to use the supporting tools that exist. I mentioned L3build and package check. And there are two other programs which I haven't used, but which are in use as I know. These are CTANify and the ctan-o-mat. Both run under Windows. Package check does not run under Windows, so maybe you will use one of the packages... one of the last two packages.

**Erik**: [04:35] From time to time some problems occur. It would be too easy if we just have the upload management. **One problem is legacy work where the author is still active, because some of the previously installed packages does not meet the actual standards.** And then we have to talk with the author to update his package so it can be seamlessly installed in TeX Live and on CTAN. And some authors say it's too much work and they will give up their package, but we try to convince them and so on. It can lead to long discussions. But in most cases the package is converted to a modern format.

**Erik**: [05:49] A similar case are packages where the author gives up or doesn't update his or her package. So it may happen that bugs are detected but not corrected. Or the TeX environment evolves, like Beamer or from the LaTeX kernel itself. And I think that was the case with the famous, often used Metropolis Beamer package, which was then replaced by the Moloch package. And **this is the hard solution we try to avoid. This is a fork.** We try to avoid it because people have to change their documents to use the new package, even if the new package has the same syntax, at least the usepackage command has to be changed.

**Erik**: [06:59] And so **it's our preferred solution that authors mark themselves or let mark themselves as inactive, and an inactive package can be taken over by every other author if it uses the LaTeX project... the LPPL.** And we are asking us if we should remind package authors of their packages, but I'm not sure. If you think we should do it, just tell me. I'm on the side who thinks the authors know their packages.

**Erik**: [07:50] Okay. Now I told a lot about we have many packages. Now I show you some numbers. As you can see, the size and the number of files and directories is steadily growing. That's the case because almost nothing is deleted. And **the size is for our times not that much. 24 gigabytes. Most mobile phones could host CTAN at the current time.** The number of authors is also growing, but as life goes on, some authors have to leave. The institutions which are also growing the number are mostly user groups and especially publishers.

**Erik**: [09:00] As I said, the TeX catalogue contains almost about TeX and it's slightly larger than the packages that are hosted on CTAN. Larger means that for some packages we have just links to their GitHub or own web page. Examples for this is JabRef. It would not make sense if Oliver and his team would update his software every time on CTAN. And some of the packages are then moved to obsolete. **Obsolete means it is not included in TeX Live, but you still find it on CTAN.** And packages get obsolete if the author wishes so, and for some very old packages which just do not run since many, many years.

**Erik**: [10:13] And for this, now I show the numbers. You see it's growing over the time. Yes, one point I wanted to emphasize: we have more small special-purpose packages. If you know my talk about the current state of CTAN from Palo Alto, the last in-present conference before the pandemic, there I said that many good solutions are hold on forums like Stack Exchange, but **now is the trend that authors, not always those who gave the answer at Stack Exchange, build packages with very small special purposes and upload it to CTAN.** This is a good trend because you don't need always the large package for your tasks. And some of these small purpose packages like TikZducks, you see the small duck, evolve to large packages. Now I use TikZducks for this nice progress bar. That means if the duck reaches the right side, the talk is finished.

**Erik**: [11:53] Okay. Here you have the numbers. Manfred collected his data from the file system, from the files we have from the catalogue, and from the Subversion commits to the file system and catalogue. And here you can see how much we have to do, how much the CTAN team has to do every year. The commits are constant almost. Updates are more and more. The reason for this is that many people use GitHub or Git and release more often as in earlier times. The number of new packages is decreasing. I have some theories about this, but it's speculation. Maybe you have a theory too. I would be... I would like your input about this number. And the lazy days are days without a commit in the Subversion repository. As you can see we don't have lazy days, and lazy doesn't mean we are lazy, but it's just no commit.

**Erik**: [13:28] So in conferences I am asked in informal discussions often questions, more often than other questions. And I compiled them here and I made notes about this. Okay. **Why do we maintain our own mirror infrastructure instead of using a commercial CDN?** And we don't use software like MirrorBrain or a commercial CDN like Cloudflare, which seems not to cost that much for a smaller network like CTAN compared to Netflix or so. And **the risks of a commercial CDN is simply that it uses black-box routing. We cannot be sure that requests are routed, and the server may be disrupted by political sanctions.** Maybe we had this in the last years between Israel and Syria, and between Russia and the Ukraine. Then we got messages from users from the Ukraine. They were redirected to Russia but could not access files there. And so we have our own redirection with exceptions of the sad cases of political disruptions. And we don't use the great software MirrorBrain because it's not maintained anymore and bugs are not corrected, so... and it's more fun to correct or to fix our own bugs than bugs from other people.

**Erik**: [15:35] Okay, next question. Yeah, it's a similar question. **Why do we run our own mail system? It's the cost, because the mailing list archives is now so large we could not provide it via a commercial provider for less money, but it's simply expensive.** It would cost us more than $100 per month. And another point is that the privacy in times where the AI companies scrape every data they can get, we don't want to offer our many gigabyte large mail conversation to Google or so, even if the company says they do not use it, they may change their mind. You had it on Reddit a year ago, Reddit said it's all private, now they sold their data. And we don't do it. We host it by ourselves.

**Erik**: [16:59] Yeah, same question, **why do we host our own name servers? It's for... yes, we can implement version control, we know when and who changed something in the name server.** And additionally, it's easier to give people access via SSH on a server than to a web page of a provider, even if it's simpler at the first look. If you use a web page, text files are better to handle.

**Erik**: [17:45] And another question, it's a question I'm happy about: **how can people join CTAN or help TeX in general?** CTAN is very time-consuming, at least for the upload management team. And so we recommend write your own packages or go on the repositories at GitHub or GitLab and try to fix issues or try to send them issues. People can work at TeX Live or MiKTeX which are on GitHub too. And I don't want to send people back, you may try to join the upload team, but I warn you, it's a long-term task for years with tasks for every day. But if you have time, just tell me after the conference or after this talk. Do you have any questions?

**Audience Member 1**: [19:18] Yeah. So I saw the new design of the website and I was just wondering if we have any plans to make it work with JavaScript-free browsers and the browsers that disable JavaScripts, or some engines like LibreJS, if we have any plans to make it compatible with such settings?

**Erik**: [19:48] Did I understand it correctly, that you ask if JavaScript has to be used or not?

**Audience Member 1**: [19:59] No, I meant if I disable JavaScripts in my browser... So the licenses are free software licenses for the new JavaScripts, yes. But sometimes the detection programs don't detect the licenses properly. Like LibreJS, for example, sometimes fails to detect free software licenses, and then the site breaks. Or if I just manually disable the JavaScripts, some sites break. They don't load properly.

**Erik**: [20:34] Oh, that's a good question. I have to forward it to Gerd, our programmer. But I will do it. And if you send me a message, I can answer to you directly.

**Audience Member 2**: [20:52] I would add to the previous question, for some... there's an accessibility issue because **for blind people JavaScript might not work well for text-based browsers like Lynx and so on.** So because the main site is in Germany and European laws, you probably want to have a text-only version without scripting and so on, which would be available for blind people.

**Erik**: [21:22] Okay. Yes, yes. I will tell this to Gerd too. And I forgot to say, **we have a programming application interface. You may still use this interface. L3build and ctan-o-mat use it too.** And it should be possible to use this application programming interface for blind people too.

**Moderator**: [21:54] Any more questions?

**Audience Member 3**: [21:58] Yes, I have. I have a suggestion and a question. The suggestion is about the inactive authors. **Most Linux distributions have a policy to contact inactive authors. They try to reach by email or otherwise. And after a certain number of tries, they are considered really inactive and somebody else can take over the package maintenance.**

**Erik**: [22:22] Okay, two types of inactive. Okay. I'll remember this too. Thank you.

**Audience Member 3**: [22:28] And the question is... CTAN is really open. **Do you get bogged down by the AI crawlers nowadays which is scraping everything?**

**Erik**: [22:38] Yes, this is a hot problem. You may have noticed that sometimes the page breaks for a few minutes. Now we have installed a monitoring software which monitors and restarts the portal. Yes, that's really a problem.

**Audience Member 3**: [22:58] And do you implement any defense mechanisms like [abuse?]?

**Erik**: [23:06] We use fail2ban and simple network exclusion from the IP addresses of the scrapers.

**Moderator**: [23:20] Quick question.

**Audience Member 4**: [23:24] Any challenges related to viruses?

**Erik**: [23:28] What, please?

**Audience Member 4**: [23:30] Viruses. Any challenges related to viruses?

**Erik**: [23:34] Viruses. No. It may be possible that people upload viruses, but it never happened until now.

**Moderator**: [23:53] Okay. Again, a great applause for Erik Braun.

**Moderator**: [24:00] So now we'll have a pause for 20 minutes. Please be back before 11 o'clock.
