Transcript: Get discovered: How to use your metadata to conquer AI search

Tim: Hello, everyone. Thank you for joining us for today’s Tech Forum session. I’m Tim Middleton, bibliographic manager and product manager at BookNet. Welcome to Get Discovered: How to Use Your Metadata to Conquer AI Search.

Before we get started, BookNet Canada acknowledges that its operations are remote and our colleagues contribute their work from the traditional territories of the Mississaugas of the Credit, the Anishnawbe, the Haudenosaunee, the Wyandot, the Mi’kmaq, the Ojibwa of Fort William First Nation, the Three Fires Confederacy of First Nations, which includes the Ojibwa, the Odawa and the Potawatomie, the Métis, as well as the unceded and ancestral territory of the Musqueam, Squamish, Tsleil-Waututh peoples, the original nations and peoples of the lands we now call Beeton, Guelph, Halifax, Thunder Bay, Toronto, Vancouver, Vaughan, and Windsor. We encourage you to visit the nativeland.ca website to learn more about the peoples whose land you’re joining from today.

Moreover, BookNet endorses the calls to action from the Truth and Reconciliation Commission of Canada, and supports an ongoing shift from gatekeeping to space-making in the book industry. The book industry has long been an industry of gatekeeping. Anyone who works at any stage of the book supply chain carries a responsibility to serve readers by publishing, promoting, and supplying works that represent the wide extent of human experiences and identities in all its complicated intersectionality.

We at BookNet are committed to working with our partners in the industry as we move towards a framework that supports space-making, which ensures that marginalised creators and professionals all have the opportunity to contribute, work, and lead.

If during the presentation you have questions, please use the Q&A panel found in the bottom menu.

Now, let me introduce our speaker, Tricia McCraney. Tricia and I have a long history together in this industry.

Tricia: We do.

Tim: Tricia works everywhere books and technology meet. She has worked in the book industry for 30 years and has a deep understanding of publisher workflows, title management, systems, and industry standards. Along with our stories, the people who tell them, publish them, and distribute them, Tricia believes metadata is our most important asset. Thank you for joining us today, Tricia.

Tricia: Thank you, Tim. I’m happy to be here. I will talk about what everyone is already doing in terms of their metadata, what we should be doing next, and talk about what we should be considering with AI now existing in our metadata landscape. So, some of the things I’ll talk about today are what’s new, what we should be thinking about now, what’s new and what will stay the same in that sense. I’ll talk about specific metadata, and give some guidance about keywords, sales points, and descriptions, and other data points you might not be populating today. And I’ll talk about evolving standards, how ONIX is evolving and how it might change with AI, and then talk about the long tail and how we can get more attention on our backlist titles, and maybe think about what libraries can do. And some of these things really came from questions that were submitted by the audience for this session. So, I’m happy to take questions at the end as well.

So, from a metadata perspective, what’s new and what will stay the same? Well, the first thought that came to my mind when thinking about this question is, well, we’ve been here before. You’ll notice that, throughout the presentation, I’ve put cover images that are sort of relevant to the theme of the slide that are representing our Canadian publishing output. So, hopefully, you’ll find one of your titles in the slides. And if not, I hope maybe we’ll sell a book or two just because the covers are here. Thanks to 49th Shelf for being a good resource in that regard.

So, as many of you know, who are on this call, the book industry is always facing some kind of crisis. It’s been reshaped roughly every decade, and each…and I’m starting…when I’m thinking about that, I’m going back to the 1970s. I wasn’t working in the industry then, but that’s sort of a good starting point if we think about how the modern book industry came to be. And each new change or shift across those decades has really one thing at the heart of it, and that is that discovery or the way that our readers find books has moved further from the physical book, from that tactile experience of holding the book in your hands, and deeper into the information that we provide or the data.

So, I’ll just do a little quick kind of walk down memory lane here. If we think about where the industry was and where things were, in…oh, hang on one second. Sorry. I’m missing a slide there. Sorry. Okay. Anyway, I will…sorry about that.

Okay. The slide that I’m missing is 1970 to kind of the ’90s. And I think what we all know happened in that period is, in the ’70s, we started with identifiers. So, ISBNs came into being. And then there was sort of a good run of independent bookstores and discovery happening in those bookstores in the ’80s. And then in the ’90s we had…towards the end of the ’90s, we had the advent of Amazon. Amazon came into being, and in Canada, we had laws in place that said that even though Amazon was not operating in Canada or didn’t have employees in Canada, that they could still continue to operate here because of the online nature. And that was an interesting decision from a Canadian perspective.

During that same time, we had Chapters come into being in Canada, and we had Indigo come into play. And towards the end of that period of time, in the sort of early 2000s, that’s when we started to see eBooks come into play. And really, a shift towards eBooks started to come into place mid-2000s. Now, we’re finally at the slide. Sorry about that. Where in 2007, Kindle launched and eBooks really began to sell. In 2009, they became real for all of the publishing community. Everybody needed to start releasing eBooks. Indigo launched what was then called Shortcovers, and we know that became Kobo. And at that time, in 2009, that’s when ONIX 3.0 was released. And that’s where we started to have better handling of digital products and formats and territorial rights from a metadata perspective.

Then we kind of dealt with eBooks for a period of time, and that was a big focus around 2013, 2014, and kind of up to the beginning of 2020. This is where Amazon, of course, really became the dominant player, and Thema came into place as a subject classification scheme that could be used internationally and not just for a single market. It replaced BIC in the UK, but the aim of Thema is to be a global subject standard. And during this time, this is when BiblioShare, CataList, and 49th Shelf — so, all of our good Canadian initiatives that I’ll call the Canadian data layer — this is where they all kind of started to flourish and grow.

And now, we come to sort of where we are today. In 2020, we had a bit of a surprise boost with the pandemic because book sales increased during that period. And we’ve had BookTok coming on the scene and influencing things that we can’t predict. Some backlist sales are surging and growing because of the influence of BookTok. And now, we have ONIX 3.1 in place. And now, today, we have AI rearing its head in our workflows.

So, all of that to say, we’ve been here before. And now, this is the new thing that we’re facing in the book industry.

What happened across that time is, like I said, discovery has really increasingly moved off the shelf and into those electronic product records. And your metadata has been machine-read for 25 years. Some time ago, and this is quite a while ago now, I can remember hearing David Caron of ECW Press say that every book that you sell is being sold out of a database, and that was probably 15 or 20 years ago. And it’s even more true today.

Today, AI is in place, and it’s…really, this is not the first time that we’ve had machines influencing the supply chain. So, the lesson for us is that what do we need to do? We need to keep on doing what we do well. And across those 25 years, the best practice has really been the same, and that is to maintain complete, comprehensive, structured data that’s consistent and accurate. And there’s a little more focus on what data we’re providing now. So, really, what we need to be thinking about is this lesson that it’s important to own your data and also supply it early. Owning it means supplying it early. If you’re the first one to supply it, that’s the best case scenario.

So, today, where we are is, as an industry, we’re really being faced with a distinction that we haven’t had to make before in terms of the data that we’re supplying. Publishers in Canada and globally are pretty comfortable with metadata now. We’ve been doing it, like I said, for 20, 25 years. And what we’re doing with metadata is now being…it has some nuance to it that it didn’t have before. So, we need to think about metadata in search and how search engines act on that metadata. That’s always been true. But now, we need to think about that data versus how it acts and is used in the supply chain. And I’ll talk about each of these things in some upcoming slides.

So, just to set the stage, there are broadly three types of metadata. And this classification here that I’m talking about really comes from our friends in the library and information science world. They love taxonomies, maybe even more than we do as publishers. But there’s descriptive metadata, structural metadata, and administrative metadata.

Descriptive metadata tells us what the thing is. So, what is our product? There’s a title, a subtitle, there are contributors. It describes subjects, categories, classifications, and the language that the book is published in, for example.

Structural metadata is really how the parts of the thing relate to each other. So, this is where we start getting into chapter-level data, page count or extent of the product, the table of contents, EPUB navigation, and even some accessibility features. Some of this data is data that we are supplying for the first time or newly supplying or recently supplying.

Administrative data is metadata that describes how the thing, your book, is managed. So, this is where we’re now…this is entirely new. So, we have the core descriptive stuff, the structural stuff that we’ve been starting to supply and are doing to some degree or starting to, and the administrative data, which is where we’re talking about rights, including for AI, how our data can be used by AI, if at all. It’s managing stuff like technical specifications. So, things like some of the accessibility features I mentioned before, but also things like the actual spec of the book, the physical spec of the book, the production information about the book.

And then the last thing, when we’re talking about administrative data, is the provenance. And that’s what I’ll talk about a little bit more throughout this presentation is where is the ownership in the details of the book. So, we’re almost going from a more rights-based and production-based view than we have before.

As you can imagine, the descriptive metadata is where we’ve put all our focus in the past. It’s kind of the attention-getter. Structural metadata, just to sum it up, is what makes our navigation and accessibility work. And the administrative stuff is where that AI piece of describing any AI attribution lives, as well as the terms of AI acting on our titles.

Okay. So, what does this mean in the real world? If we think about metadata and search for our end user, our readers, the biggest thing is that, now, the search box is an answer engine. So, the people who we want to buy our books, who were typically searching for them in retailer search engines before, are now chatting with search engines. So, the behavior has changed. If we talk about SEO versus AI in search engines, we need to start thinking about putting full sentences and phrases in our keywords, not just individual words. And I know phrasing has been a thing before, but it’s actually more of a question that we’re answering than just putting in data that helps someone get to the end result.

Matching and ranking that happened in SEO — so, where your search terms returned a list of links or possible results — has now moved to semantic retrieval. And that means not just matching and ranking, but AI is making a decision about what to display back to someone using a search engine. So, instead of that list of links, the users, us as readers or our end users, are getting a synthesised answer.

So, practically, what that means is that if someone is asking what should I read next or what would my relative who’s a 12-year-old who likes graphic novels like to read next, they are now typing that full question into a search bar, not just typing in the keywords that might form that question. And the search functions on retail sites now have AI layers acting on them. And so, the results that come back are not always just a synthesised answer, but you might have seen the experience. I haven’t seen this on our major book retailers yet, but you might have had the experience in other sites where results come back and they sometimes prompt you to refine the results that come back. They ask you more, just like if you’re chatting with an LLM, Claude, or one of your language models. Just in the same way, it might just ask you more questions to refine those results. So, your data needs to be able to respond to these questions and not just serve up individual data points like categories and classifications.

Now, what happens in search or what’s happening in search right now is a little bit different than what’s happening in the supply chain. For suppliers — so, those recipients of your metadata — their systems are really looking at everything. So, suppliers will mine your data, and that’s always been the case. They’re already looking at your data and extracting whatever you send, even if we can’t see it pushed through onto a product page. They’re using that data in meaningful ways. But the way that categorisation is happening with suppliers is shifting. So, they now want — and in some cases, need — to know more about how your books are made.

And what’s important for us to know as suppliers of metadata is that any records that don’t have enough data, so that are sparse or thin, can get interpreted poorly. They basically read badly in AI, and then we have the risk of AI hallucinating or making something up. And the same is true if data is contradictory.

So, in this current area, in this world that we live in now, your metadata is being interpreted and not just displayed. That’s kind of the takeaway here. Supply chain systems used to take your data and pass it through, and you would kind of evaluate what you sent based on what you saw on individual product pages and different sites. But now, that data is being evaluated against itself, your own records and other records. And the systems and suppliers are making decisions about what the book is and who it’s for.

So, what does this mean in terms of standards and what we can do? In terms of AI, there are some considerations. And this kind of goes right to the heart of the AI question. ONIX already carries an ability and a mechanism for communicating AI involvement, as well as a publisher’s terms for text and data mining, which we currently refer to and people describe as just TDM. And I’ll talk about that. I’ll talk about the AI involvement in a later slide, but the terms for text and data mining, there’s a little snippet here in the right that shows you how to describe the fact that text and data mining is prohibited in your books.

So, we have an EPUB usage type and an EPUB usage status here. This is within the EPUB usage constraint composite in ONIX. And when we see that EPUB usage type of 11, that indicates text and data mining. And the status of 03 tells us that it’s prohibited. And these are from ONIX code lists 145 and 146. And what this tells us is this has been in place in ONIX for some time. So, the ONIX standard was kind of ready and waiting. It had already moved to address AI before most of the rest of the industry did, before systems and suppliers and publishers did.

And this type of information about copyright and AI text and data mining, this is…policy around these things is being developed in publishing houses and in companies in every major market where information is shared. And that, of course, includes Canada. And BookNet Canada has a great blog post about this that I would refer you to. And I think Nataly will put a link in the chat.

Okay. So, that’s kind of what’s happening today. And what nobody actually knows — this title “The Unnamable” from Nimbus was just great to go with this slide — is how much book discovery flows through AI channels now. I’ll just say it again, we don’t know how much book discovery flows through AI channels now. That’s a little haiku. It’s fun making that for this slide.

So, this is a question that we need to pursue, and it’s worth asking the recipients of your data what they are doing in that regard. Are they using AI in their backend systems? And if so, what are they doing with it? And certainly, that’s a question that BookNet has been pursuing. They have a paper that’s coming out in the fall that was conducted together with BISG on what’s happening with AI in the industry. I won’t say anything more than that, but it’s coming in the fall. And of course, it’s of interest to all of us.

Okay. So, that’s kind of where we are today. It brings us up to date with the most current stuff, but what’s really new and what should we do? So, let’s think about some actions and what we as people who create metadata, circulate it, and publish books can do.

The things that are new now really centre around the source material of…so, the source for metadata. Oh, sorry. I have something strange there. Ignore that. The source material, the provenance, and potential AI hallucination.

So, in terms of source data, the words that make up our data are source material for discovery. And that means that they are source material that, like I said, can be considered and evaluated and had decisions made upon by others. And this means that we need to be really clear with our data, we need to be concise, and leave less room for ambiguity.

So, when it comes to the details, no previous shift in the industry, like I talked about from the ’70s until today, has really asked for declarations about how our books are being made. And that’s being asked for now because the way our books can be made has changed. There are some things that fall in this category of the details being different. The devil is always in the details, but there are some things that come up around other industry standards, and I’ll mention those in another slide. But really, what we need to be thinking about right now, newly, is the source material for discovery, which is our data, the details that we’re providing in there, and there are new details to provide, and consistency.

So, we used to think about flavors of ONIX as being a thing and different metadata for different audiences used to be something interesting and something to pursue and something worth doing. Today, with AI in the picture and on the scene, different flavors of ONIX and different flavors of our data might be confusing to AI. And that’s where we get that kind of potential AI hallucination scenario.

So, in terms of specific metadata guidance, some of you submitted some questions and I’m attempting to answer those, but really what we’re to do today is get even more excited about your metadata. So, if you have metadata fields that you consider optional or even that are considered optional or deemed optional in our metadata schemes — so, in guidance from your recipients or suppliers, distributors, or even from BookNet Canada and others — you should really consider, and we want to think about supplying as much metadata as we can. And looking at those optional fields is really the first place to go.

So, this means descriptive copy that goes beyond just the main description. And so, naturally, you’re thinking, I’m already creating a main description and a short description, what else do I need to do? Table of contents is one that is often overlooked and always good to do. Start supplying subject codes across multiple schemes, and maybe you’re already doing BISAC and Thema, but you might consider including regional themes and merchandising themes.

Contributor biographies for all roles, so, not just your primary author and maybe an illustrator, but any contributor on the title that has a public-facing role, consider having a contributor biography for them if you don’t, and within that biography, including facts. So, things like affiliations, things that are concrete bits of information that an AI model might use.

Other things you can include if you’re not doing it already are identifiers other than the holy grail, the ISBN. So, that means things like DOIs, digital object identifiers, if you have them for your digital products, things like Dewey decimal codes or LOC codes if you’re selling into the U.S., things like identifiers for your names or even identifiers for some of your series and collections.

Excerpts. Excerpts and reading samples are always great to include if you’re not doing it already. Of course, that depends on what you have the right to share, but those are really great things to start doing if you’re not already. And then a strong focus on series name and order. And this is one where I really have to say as not just somebody who is 101 years old in metadata years and loves metadata, but somebody who regularly buys books and looks for books for kids. Series and collection information is one of the places where many of us are not doing it good enough, and that includes the places where the information gets displayed, of course.

And what I wanted to say is, in addition to making your own series information strong and accurate, a series can be more than just a list of books in the typical kind of numbered order that we think of, one, two, three, four. But a series, we can also think of a series as a collection of linked titles. And linking your titles together in meaningful and even creative ways can be really helpful with discovery.

So, you’re probably already creating your own collections that group titles together in some way based on your own marketing efforts or promotional efforts or even your own subject classifications. And this is one thing that you can supply in ONIX. It allows for proprietary collections. So, collections is that term that we use in ONIX 3.0 and 3.1 to describe what we would traditionally think of as a series. And proprietary ones can be sent. Here’s an example. And these are a way of, like I said, linking your titles together so that a machine is reading that information, and it’s providing information that maybe we can’t see on the end result, which is the Amazon or Indigo or another distributor’s product page or web page. But it informs what you as a publisher know about the book.

So, this just gives a snippet of what ONIX looks like if you’re communicating your own collection in ONIX with a code and a title. So, in this case, we can see we have a collection that is called a publisher collection. So, that’s just a collection type and a name that tells the recipient who this collection belongs to. And the interesting part down below is that we’re calling it Summer Reads 2026. So, that’s the name of this collection. It’s not the name of the title itself, but that title composite here is our collection.

Now, one cautionary note about this is that it can require a bit of upkeep. This is a deliberate example of something that is seasonal. And you might want to update this or remove this collection from your metadata after summer 2026 or maybe at the end of 2026 after a certain period of time passes. Okay. So, I can’t stress linking your titles together in kind of new and interesting ways enough. And of course, that collections area is one way to do it.

Another thing that, of course, is really important and that is changing in terms of what we do with this metadata is keywords. And one of the questions that came up before this session was we have one place where we can provide keywords in our metadata, and we have many recipients. What should we do?

So, the way that you are approaching your keywords is now changing. Like I said, we’re using phrases, we’re answering questions instead of using single words. But additionally, you want to put your most important keywords first because different recipients have different limits for what they ingest with keywords. And that’s always been true, but it seems to be a little more true with AI. AI tends to prioritise the first couple of sentences of descriptive copy, and really tends to emphasise the first, the initial keywords.

So, you’re writing your keywords now, not for everybody who might take as many as they take, but for the recipient with the most strict or the tightest limit. So, if you have a recipient who only takes 20 keywords, that’s what you’re aiming for while everybody else takes more. All of your keywords and phrases, if you’re working with that short limit, should count.

So, some things that you don’t want to do are you don’t want to use bestseller or any deliberate misspellings because there is a level of fact checking that happens with AI. If it’s not actually a bestseller, it can cause confusion. And importantly, I’m sure we’ve said this before in metadata circles, but your keywords, whatever you enter in your keywords, should not exist anywhere else in your metadata. So, if information can already be held and communicated and ingested in other places within your metadata record, don’t put it in your keywords.

I would make a small caveat about series information because, in my experience, series information isn’t being ingested really well and used really well. So, that’s a place where you might repeat a series name. That seems like a fair caveat.

And one other really important thing is don’t do keyword stuffing. Keyword stuffing is…I’m sure everybody knows. But just to be clear, keyword stuffing is when you have a title of the book or a subtitle, and you’re stuffing or adding keywords into that data specifically for discovery purposes.

So, we might see, for example, the title of the book on this slide is “The Many.” And you might see…I’m not saying that the publisher of this book does this, but you might see a publisher, who is keyword stuffing, send “The Many” as the title, and from the bestselling author of “Sleeping Giants,” which is in the text on the front of the book within the title itself, because then they’ve got the word bestselling in there.

So, there are a couple of reasons not to do keyword stuffing, and one is it does create confusion. And like I said, we want our data structured properly, and we want our data in the right place. But beyond that, my kind of personal take is that it’s just sort of off-putting. It’s kind of patronising. If I’m looking at results and I see a book where there are tons of keywords stuffed into it, it just feels like I don’t know what I’m looking for or I can’t…I find it personally just kind of off-putting. I’m sure many of you do, too.

Okay. And so, some of that stuff is information that we already know, but it kind of becomes even more important now with AI on the scene. And I wanted to also bring to the front a few more things to consider when you’re adding to your metadata. And this is relevant to keywords, but also, this can be relevant to wherever it makes sense to add this data. So, whether it’s in your copy or your metadata or specific places.

Other things to add and think about that you may not be doing, if you’re adjusting your metadata, are subgenre subjects and categories. So, there might not be enough subgenres and categories in the controlled vocabulary schemes in BISAC and Thema. You may want to add a proprietary subject code to your ONIX. So, in the way that BISAC and Thema have their own category values and in the way that we are able to provide those proprietary collections, you can set up and send your own proprietary subjects and categories. So, this is the kind of thing that you might do if your website has a subject scheme or a subject tree that’s different from those industry standard schemes. You might decide to include that in your ONIX because it’s more descriptive of your own mission and what the press does or what you do as a publisher.

You want to think about adding tropes. So, this is, of course, seeing a lot of traction with all of our romantasy books. These are the enemies to lovers and all of these kind of funny things that come up, but they’re a very significant and useful thing to add to your books where it makes sense.

Mood and tone are something that maybe we aren’t always including in our keywords. Place setting, including regional specificity. So, of course, we can do some of that in our BISACs and our regional codes and in Thema, but if you need to be even more specific, certainly do that in your keywords and descriptive data where it makes sense.

Now, here are a couple that I’ve seen being used broadly by some of our customers. Where I work at Virtusales, we have a title management system called BiblioSuite. And so, I’m working with publishers to implement their workflows in the system, and that includes metadata. And what I see is a lot of publishers adding holiday and occasion and gifting language into their keywords, their metadata, or even into a proprietary subject scheme like that that might hold something like books for Easter or books for Mother’s Day, things like that. And you can go, obviously, more granular than that.

One thing that I’m seeing also is publishers adding relevant current events to their keywords and metadata. Really, what I’m seeing is publishers adding current events to their metadata. So, I’ve added relevant. It should be relevant. If there’s something going on in the world and it’s making headlines, there are publishers who are just adding keywords to every book in their catalogue to see what happens, or with the theory or the gamble that this is going to increase visibility and discoverability. So, I would say, current events or things that are topical and very now are good to add, but make sure you’re refreshing that information.

And finally, believe it or not, I’m seeing a lot of publishers adding emojis to their metadata. I tried to pick a neutral one to show here on the screen. It could have been mind-blown or any other number of emojis easily. But in this case, we are seeing in descriptive copy. So, not in keywords, but in things like your descriptions and short descriptions and kind of promotional headlines and things like that, emojis being added. So, that’s one to consider. I’m not recommending it or un-recommending it, but that’s something that’s happening.

Okay. And really, when you’re going through that exercise, the important thing to really do is to experiment. I talked about the importance of renewing your metadata. And really, when we think about it, metadata is a renewable resource. The person power to create and refresh the metadata isn’t always a renewable resource. So, of course, that’s a challenge. But maintaining your own guidelines for writing keywords and copy, and renewing those is just as important as actually creating the metadata and updating it.

So, as a publisher or an imprint within a publisher, you want to know what you’re trying to do with your metadata and how it fits kind of the overall mission or the overall ideal of the press or the title itself. And those guidelines are really something that should be revisited. If you’re doing metadata today the same way that you did it 5 or 10 years ago, it’s a good time to think about what do we do when we write metadata? What are our internal guidelines? Do we have internal guidelines? And should those change and be updated?

So, if you are able to update your metadata frequently, of course, that’s ideal. But if you can’t do that frequently, experiment, choose a subset of your data. Backlist is always good to experiment with. And track what happens when you update it, when you implement some of these changes, and when you send that data out into the world. As I said, we really don’t know how much AI is acting on book discovery right now or how much book discovery and metadata is flowing through AI channels.

Tim mentioned gatekeeping and publishing at the beginning of this session, and I want to say that Canadian publishers tend to be more forthcoming and willing to share. So, if you are experimenting and trying new things, share what you know if you feel comfortable. I mean, everybody wants to know what we’re doing and why we’re doing it and what’s working. And it takes some effort to track those things and figure out, is a book that I have revived with new metadata actually seeing an impact in sales, or is it getting in front of more people? So, please, let’s kiss and tell, the book says.

We can talk about evolving standards and how ONIX and other industry standards kind of help us cope with what’s going on. I talked about that text and data mining mechanism in ONIX earlier, but another thing that we can talk about is declaring AI involvement. And ONIX provides a way of communicating the involvement of AI in your titles. And so, that’s where contributors are assisted by AI. This is where ONIX allows us to say there’s an unnamed person involved.

So, looking at the ONIX snippet on the right, that example…this is the worked example from EDItEUR. And we can see two contributors here. The first one, James Green, we see A01, that’s our author. The second one, Fiona Brown, has a contributor role code of A12. And Fiona is the illustrator in this case. The third contributor here has a contributor role of Z01. And that’s where we’re really saying assisted by and 09 is an unnamed person. That’s our indication of AI there.

And the sequence that we see here is really important because this is saying James Green is a person who’s an author. There was no AI involved in James’s contribution. Fiona Brown is the illustrator, and because we see the assisted by an unnamed person immediately following her contributor record here in sequence, this is telling us that the illustrator was assisted by AI.

So, ONIX has already evolved to the point of allowing us to communicate this type of thing within the standard. And there are a couple of other things that I think might happen in terms of how it might evolve. It’s really expanding now to hold more information, like I said, about the books that we’re creating. And the biggest thing that I think is going to happen is that ONIX is no longer going to be just our marketing message.

So, that’s really kind of what it’s been up until recently. It’s been a container and a place to hold all of the information that we want to go out into the world and sell our book. It’s going to continue to evolve to hold more and more information that’s not directly related to marketing. So, today, it can contain not only information about that AI involvement or the permissions about text and data mining, but also about accessibility and print-on-demand specifications. It can hold information about regulations like the EUDR, the European Union Deforestation Regulations. It can hold information about global product safety, the GPSR stuff, so, contacts there.

And with this information now being expected to communicate in ONIX, it’s becoming clear that ONIX is, like I said, not just that marketing message, but it’s our common language. So, anybody you’re sending data to might expect it to be held in ONIX. And ONIX is really…it’s there for us. We can contribute to it. So, it evolves with what the industry and the end users of that message want.

And I think, on that note, the thing to ponder or to think about is, what’s the difference between my metadata and my? So, ONIX is now going to hold a lot of that kind of rights-related or provenance-related data, and that will continue to evolve, I think.

So, your metadata is kind of what we think of, traditionally, as all that marketing data. And it’s the stuff that we want to go out into the world and circulate everywhere and anywhere it can go, as widely as possible. And ONIX is really, like I said, the carrier or the container for that metadata. And that means that when you send an ONIX message or a product record out into the world, there’s an implied license to ingest that data and now to mine it for anybody who receives it. And that means that the metadata you put in ONIX can always be ingested and mined, even if you opt out of text and data mining for the content of a book. So, your metadata should only hold data in ONIX that you want to be available and used in some way by the end recipient.

Your content, well, that’s the lifeblood of publishers. And that’s where you want that information and that data to be used on your terms and according to the rights that you have. This is where you want to include those constraints in ONIX and on your website, like saying we do not allow this data to be mined. And consistency with what you communicate around rights for your content is really important because it limits your exposure.

There was a question about how do we communicate that our books are written by authors and not by AI? So, our books are written by humans and not by AI, and how do we communicate that? There isn’t a good and substantial answer for this question, but the short answer is, if you want to say our books are not written by AI, you don’t. It’s a deliberate thing that EDItEUR has done to say where AI is used and not to say there’s no use of AI.

One thing to think about here as I was pondering this question is, as publishers, as people working in the book industry, can we really say confidently that no AI was used in our books? If we think about things like the cover, are we sure that AI wasn’t used in the cover? Are we sure that it wasn’t used in copy editing? Are we sure that it wasn’t used to create an index? Maybe. We hope so, but that’s definitely one thing to think about. How confident are we that no AI was used at any point in the process? And when it comes to ONIX, that vocabulary that they’ve developed that we looked at is really about declaring where AI is used and not the absence of AI. Like I said, it was a deliberate choice by EDItEUR.

So, now, I want to talk about what do we do for backlist titles and where do libraries fit in? We’re getting a little bit short on time, and there’s not a lot more here to cover. But backlist titles are a great place to start in terms of where you experiment with the metadata and start adding stuff because, of course, they have usually the least complete metadata.

So, you want to be running compliance reports for metadata completeness, and you want to look not just for whether or not data is missing, but also for incomplete metadata and stuff that’s wrong or incorrect. So, in this sense, I’m saying, work the gaps and not the individual titles. And what are some of the gaps? Any awards that might have been won since publication.

Links, so, related products. And like I talked about, series and contributors because those links are really what will bring your titles along when new books come out and surface them and make them more discoverable.

Any current hooks or updated language. So, an example of updated language is something…we don’t say homeless anymore, we say unhoused. And things with phrases and language that come up, you can update new phrases and new things that are being used just generally across social media, too.

Anniversaries. Everybody loves the comeback. So, start marketing the 10th anniversary of some of your more popular books. That’s always a great place to start.

Somebody asked, what can libraries do? And my short answer here is, really, libraries are already doing a lot of great things that publishers might learn from, too. They have a lot of authority and discipline applied to their metadata already. So, they have consistent forms of authors and contributors names, same for series. And they use identifiers for those things so that we know what we know is who and what we know it is. They also, though, could be looking towards curation with things like BookNet Canada’s Loan Stars programme.

To summarise, I did want to say something about kind of what the takeaways are. One is fill the fields you’ve been skipping, not just for front list, but also for your backlist titles. And get curious about what happens when you do that. So, track it, start to figure out the impact. Keywords contain information the rest of your data doesn’t and nothing else. And start creating them so that they’re answering questions. Say what’s true about the provenance of your books in the structured fields for that data.

And finally, somebody asked something about sales points that I didn’t answer in the slides. But sales points now need to be written with the assumption that a salesperson is not in the room supplementing them or adding to them. So, you want to write sales points so that a machine is getting all of the information that you might normally communicate in a sales meeting.

Two minutes left. Thank you. I’m happy to answer questions afterwards by email or through BookNet, as I realised I didn’t leave a great amount of time for it.

Tim: Thanks, Tricia. Thank you for that awesome presentation. And yes, you’re right. We do have a few questions.

Tricia: Thank you, Tim. Okay.

Tim: We do have some time. Lots of questions came in, lots of people are thinking about this, absolutely. It’s fast moving. And so, we need all the info we can get. So, I’ll just…we have some prioritised ones. So, the first question is, if you prohibit text and data mining in the EPUB coding, does it also restrict the book from appearing in search results because the AI is not allowed to scan the material for an answer?

Tricia: Oh, good question. No, it will not. That’s the difference between the metadata and the content. So, it will not restrict the book from being returned in search results, I’m happy to say. Yes.

Tim: Yeah, very important. Okay. Moving on this next question. When consistency talks about one set of metadata, does it suggest one set of data regardless of context? Example, catalogue versus retail site, particularly descriptive copy, I think.

Tricia: That’s a good question. I think it doesn’t have to be the same descriptive copy that you put in your catalogue versus what you send out in your metadata to a retailer. But you want to make sure that the message doesn’t contradict itself. So, the description…of course, it makes perfect sense to have a catalogue description that’s nuanced and a little bit different than your main description that goes to the retailer. You want them to work together and reflect the same picture of your book and not contradict each other. And also, you want to make sure that both of those things are populated. You can send both of those descriptive copy types, do both and do all of it.

Tim: Yeah. I had a thought that has left me, but that’s fine. I guess, actually, what I’m thinking of — and you did bring it up in your presentation quite often — is to think about your recipient, really. And keywords, perfect example. How many keywords does somebody take, how many characters can a title contain? So, anyways, that’s still so relevant. You need to understand what your recipients of your data can handle as well as thinking about your different properties like catalogue versus a retail site. Just throw in my two cents whenever I can.

Okay. Maybe a couple more questions. Should I be including duplicating my BISAC codes and keywords?

Tricia: No.

Tim: Thank you.

Tricia: Right. So, this is a good…I mean, everybody does it. So, it’s not a…everybody has done it, let’s say. But this is a good example of where a BISAC subject code has a prescribed place where it lives in ONIX, and that’s in the subject composite as a BISAC code. Your keywords are an opportunity for you to provide information that you’re not providing anywhere else. So, absolutely do not duplicate it. Do not duplicate the title or the author or any of that stuff. It’s really an opportunity to say things that haven’t been said elsewhere.

Tim: Yes. And plus, you’re keeping those keywords fresh.

Tricia: Yeah. You’re going to change them. That’s a good point, Tim. And the subject classifications don’t usually change very often. Yeah.

Tim: And then I think maybe we just have time for this. We’ll see. About the disclosure of the use of AI. Now, I know EDItEUR is really trying to…they’re trying to deal with this. They’re linking out to certifications, human-made and things like that. Nobody is completely satisfied with the solutions yet. So, this person…about the disclosure of the use of AI, shouldn’t it be explicit that it is not unnamed persons, but shouldn’t it be we used AI?

Tricia: I guess it’s about the specificity. ONIX, like I said, it’s our common language. And it allows us to say where AI has been introduced. And so, unnamed person is an option within the contributor composite in a place of saying we used AI. There are a few other options in there like synthesised voices and things like that if it’s an audiobook. But this is the kind of feedback that you can give to BookNet Canada to take to EDItEUR because, like I said, the ONIX standard is written for us, and it addresses our questions and our concerns and our real life problems.

Tim: Yeah. There’s EDItEUR or…I guess they seem to be really dealing with de minimis. So, if you’re using Grammarly for something or…in these cases, they’re struggling with how much do you disclose or depending on how much…

Tricia: How much should you, yeah.

Tim: Yeah. And then I think, for maybe our final question because…two quick ones. Is a series also titles published under a specific imprint which are all books of a similar theme?

Tricia: Sorry. You got to read that to me again.

Tim: Yeah. I think the answer is no. But is a series also titles published under a specific imprint?

Tricia: Oh. I see. I see. So, could a collection be an imprint? Yeah. It’s a good question. So, should I create a proprietary publisher collection of books under an imprint? And it’s a good line of thinking in terms of organising your books and what you might do with them if you’re presenting them together. But an imprint already has a place where that data lives in ONIX. And in fact, it’s a brand. So, that’s a little bit stronger than a series or a collection. A collection brings together things that the end user might not expect to be grouped together. So, for example, you might bring together a bunch of children’s books on a certain theme and call those your publisher’s collection, whatever, kids book theme. I can’t think of one right now, but that’s a better example. Yeah.

Tim: Yeah. And then for our last question, because it was your last point, I think. What fields in ONIX are best for anniversaries when it’s not a new edition?

Tricia: Good question. So, this is probably where you want to put that 10-year anniversary, even if it’s not a new edition, in the first line of your description, for example, or as a string in your keywords. It’s something where you want it to be kind of upfront and promoted. And it’s going to trigger an update in the record. So, those are two really main places where we know changes will happen. If you send a new description, the retailer is going to change it. If you send new keywords, the retailer is going to act on those. So, it could be your promotional headline or your description would be good places.

Tim: Awesome. I think we got a lot of the questions answered. It’s 3:07. We’ve kept people for pretty long. And we just want to thank you so much, Tricia. Thank you for joining us today. And before we go, we’d love it if you could provide feedback on this session. We’ll drop a link to the survey in the chat. Please take a couple of minutes to fill it out. We’ll also be posting a recording of this session, and we will be emailing you a link to it as soon as it’s available. To our attendees, we invite you to join us for the second part of this mini-series Get Discovered: Adapt Your Marketing for the AI Search Era, scheduled for September 17th. You can find information about all upcoming events and recordings of previous sessions on our website bnctechforum.ca. Lastly, we’d like to thank the Department of Canadian Heritage for their support through the Canada Book Fund. And thanks to you all for attending. Thank you, Tricia.

Tricia: Thank you, Tim. And thanks to our interpreters.

Tim: Yes.

Scroll to Top