From b87f342e9bdb9c902f0d4c6ca78eac92a2638ca1 Mon Sep 17 00:00:00 2001 From: Andreas Fehlner Date: Wed, 19 Aug 2026 18:01:33 +0200 Subject: [PATCH] fix: correct markdown and table formatting in ML-BOM guide Fixes a malformed table row, a missing list-item separator, an invalid HTML closing tag, and a broken markdown link in a table cell. Signed-off-by: Andreas Fehlner --- ML-BOM/en/0x22-Design-Model-Card-Parameters.md | 4 ++-- ML-BOM/en/0x24-Design-Model-Card-Considerations.md | 2 +- ML-BOM/en/0x92-Appendix-EU-AI-Act-mappings.md | 4 ++-- 3 files changed, 5 insertions(+), 5 deletions(-) diff --git a/ML-BOM/en/0x22-Design-Model-Card-Parameters.md b/ML-BOM/en/0x22-Design-Model-Card-Parameters.md index 4a682bc2..407d2742 100644 --- a/ML-BOM/en/0x22-Design-Model-Card-Parameters.md +++ b/ML-BOM/en/0x22-Design-Model-Card-Parameters.md @@ -28,7 +28,7 @@ Describes the general learning approach used to train the model. Currently, the | Type | Description | |---|---| | **supervised** | Supervised machine learning involves training an algorithm on labeled data to predict or classify new data based on the patterns learned from the labeled examples. | -| **unsupervised** | Unsupervised machine learning involves training algorithms on unlabeled data to discover patterns, structures, or relationships without explicit guidance, allowing the model to identify inherent structures or clusters within the data. +| **unsupervised** | Unsupervised machine learning involves training algorithms on unlabeled data to discover patterns, structures, or relationships without explicit guidance, allowing the model to identify inherent structures or clusters within the data. | | **reinforcement-learning** | Reinforcement learning is a type of machine learning where an agent learns to make decisions by interacting with an environment to maximize cumulative rewards, through trial and error. | | **semi-supervised** | Semi-supervised machine learning utilizes a combination of labeled and unlabeled data during training to improve model performance, leveraging the benefits of both supervised and unsupervised learning techniques. | | **self-supervised** | Self-supervised machine learning involves training models to predict parts of the input data from other parts of the same data, without requiring external labels, enabling learning from large amounts of unlabeled data. | @@ -59,7 +59,7 @@ Some examples of commonly referenced neural network (NN) architecture families i * **Convolutional Neural Network (CNN)** - an architecture designed to process efficiently detect patterns (like edges, shapes, and textures) to typically classify or analyze visual (videos/image) or auditory data, but can applied to text analysis, behavioral patterns and more. * **Recurrent Neural Network (RNN)** - an architecture designed for processing sequential data like text, speech, and time series, typically used for tasks where order matters, such as language translation, speech recognition and time-series forecasting. * **Long Short-Term Memory (LSTM)** - a specialized variant of a Recurrent Neural Network (RNN) architecture designed specifically to overcome the limitations of traditional RNNs in learning long-term dependencies. -* **Gated Recurrent Units (GRUs)** a specialized variant of a Recurrent Neural Network (RNN) architecture designed to overcome challenges like the vanishing gradient problem and enhance the modeling of long-term dependencies in sequential datasets. +* **Gated Recurrent Units (GRUs)** - a specialized variant of a Recurrent Neural Network (RNN) architecture designed to overcome challenges like the vanishing gradient problem and enhance the modeling of long-term dependencies in sequential datasets. * **Generative Adversarial Networks (GANs)** - an architecture used to train two neural networks, a *Generator* and a *Discriminator*, to compete against each other to generate more authentic data from a starting training dataset. The Generator tries to fool the Discriminator by creating fake data, while the Discriminator tries to identify fakes, leading to continuous improvement in data quality. Again, the list above represents architecture families commonly referenced in research to establish an understanding of general model design; however, the architectural landscape continues to grow as researchers specialize and optimize for different use cases, goals, and datasets. diff --git a/ML-BOM/en/0x24-Design-Model-Card-Considerations.md b/ML-BOM/en/0x24-Design-Model-Card-Considerations.md index 8c485b7b..d693600c 100644 --- a/ML-BOM/en/0x24-Design-Model-Card-Considerations.md +++ b/ML-BOM/en/0x24-Design-Model-Card-Considerations.md @@ -258,7 +258,7 @@ Each "consumption" entry consists of the following, which are explained in more | Value | Description | |---|---| | **design** | A model design including problem framing, goal definition and algorithm selection.| - | **data-collection** |Model data acquisition including search, selection and transfer.| + | **data-collection** | Model data acquisition including search, selection and transfer. | | **data-preparation** | Model data preparation including data cleaning, labeling and conversion. | | **training** | Model building, training and generalized tuning. | | **fine-tuning** | Refining a trained model to produce desired outputs for a given problem space. | diff --git a/ML-BOM/en/0x92-Appendix-EU-AI-Act-mappings.md b/ML-BOM/en/0x92-Appendix-EU-AI-Act-mappings.md index 498c90c2..b2a9c5a4 100644 --- a/ML-BOM/en/0x92-Appendix-EU-AI-Act-mappings.md +++ b/ML-BOM/en/0x92-Appendix-EU-AI-Act-mappings.md @@ -127,11 +127,11 @@ Subsections under Section 2, _"Lists of data sources"_, require similar informat | 1.3.(ii) | Training data size | The CycloneDX component can be used to describe a training dataset with any level of detail required. In general, this section describes the general method on how to declare public and private datasets:
• [Declaring datasets](0x22-Design-Model-Card-Parameters.md#declaring-datasets)

Additionally, other types of information about each component dataset can be provided via various fields such as:
• _pedigree_`
• _external references_ to documentation
• _properties_ for customized for tagging information to domain-specific requirements such as data sizes. | Dataset component(s):
• [component](https://cyclonedx.org/docs/1.7/json/#components)
   ▪ [pedigree](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_pedigree)
   ▪ [externalReferences](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_externalReferences)
   ▪ [properties](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_properties)

**Note**: _Ideally, each dataset would have its own independent Bill-of-Materials that fully described the details of its design, dependencies (i.e., data sources) and manufacturing which could be referenced by the AI/ML BOM._ | | 1.3.(iii) | Types of content | A discrete description of the types of content used to train a model would be provided as CycloneDX data components.
• [Declaring datasets](0x22-Design-Model-Card-Parameters.md#declaring-datasets)

Additional content information can be provided via external documentation and referenced in the model's component declaration.
• [Providing links to papers & articles](0x22-Design-Model-Card-Parameters.md#providing-links-to-papers--articles) | Dataset components, their descriptions and external references to documentation:
• [component.](https://cyclonedx.org/docs/1.7/json/#components)
   ▪ [type](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_type): `"data"`
   ▪ [description](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_description)
   ▪ etc.

Model component's external references:
• [metadata.component.externalReferences](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_externalReferences) | | 2 | List of data sources *(information about specific sources of data used to train the general-purpose AI model)* | *See details in subsections below* | N/A | -| 2.1 | Publicly available datasets | Each _public_ dataset used to train a model would be provided as a CycloneDX data component.
• [Declaring datasets](0x22-Design-Model-Card-Parameters.md#declaring-datasets)
   ▪ [Datasets as in-line information](0x22-Design-Model-Card-Parameters.md#datasets-as-in-line-information)
   ▪ [Datasets as data component references](#datasets-as-data-component-references) | Dataset component(s):
• [component](https://cyclonedx.org/docs/1.7/json/#components)
   ▪ [type]: `"data"`
   ▪ [name](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_name)
   ▪ [description](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_description)
   ▪ [pedigree](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_pedigree)
   ▪ [externalReferences](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_externalReferences)
   ▪ [properties](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_properties)
   ▪ etc.

**Note**: _Ideally, each public dataset would have its own independent Bill-of-Materials that fully described the details of its design, dependencies (i.e., data sources) and manufacturing which could be referenced by the AI/ML BOM._ | +| 2.1 | Publicly available datasets | Each _public_ dataset used to train a model would be provided as a CycloneDX data component.
• [Declaring datasets](0x22-Design-Model-Card-Parameters.md#declaring-datasets)
   ▪ [Datasets as in-line information](0x22-Design-Model-Card-Parameters.md#datasets-as-in-line-information)
   ▪ [Datasets as data component references](#datasets-as-data-component-references) | Dataset component(s):
• [component](https://cyclonedx.org/docs/1.7/json/#components)
   ▪ [type](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_type): `"data"`
   ▪ [name](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_name)
   ▪ [description](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_description)
   ▪ [pedigree](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_pedigree)
   ▪ [externalReferences](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_externalReferences)
   ▪ [properties](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_properties)
   ▪ etc.

**Note**: _Ideally, each public dataset would have its own independent Bill-of-Materials that fully described the details of its design, dependencies (i.e., data sources) and manufacturing which could be referenced by the AI/ML BOM._ | | 2.2 | Private non-publicly available datasets obtained from third parties | Private dataset information would be provided similarly to public datasets.

See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ | | 2.2.1 | Datasets commercially licensed by rightsholders or their representatives | Commercial dataset information would be provided similarly to public datasets.

See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ | | 2.2.1.(i) | concluded transactional commercial licensing agreement (modalities covered by license) | License information would be provided in the CycloneDX data component:
• [Describing models as components](0x20-Design-Model-Component-Metadata.md#describing-models-as-components)
   ▪ _[Example: Declaring an ML model in an ML-BOM](0x20-Design-Model-Component-Metadata.md#example-declaring-an-ml-model-in-an-ml-bom)_ which uses the CycloneDX `license` object. | CycloneDX provides multiple, robust options for recording license information:
• [metadata.licenses](https://cyclonedx.org/docs/1.7/json/#metadata_licenses)

**Note**: *modality-specific licensing may have considerations in future CycloneDX versions.* | -| 2.2.2 | Private datasets obtained from other third parties | Third-party, private dataset information would be provided similarly to public datasets.
See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ | +| 2.2.2 | Private datasets obtained from other third parties | Third-party, private dataset information would be provided similarly to public datasets.

See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ | | 2.2.2.(i) | Specify the modality(ies) of the content covered by the datasets concerned. | Model data component modalities are declared in the same way as for the model component itself:
• [Declaring a model's modalities](0x40-Design-Additional-Model-Information.md#declaring-a-models-modalities) | Data component modalities as properties:
• [component.properties](https://cyclonedx.org/docs/1.7/json/#metadata_tools_oneOf_i0_components_items_properties)

**Note**: _Utilizes property values defined in the the [CycloneDX Property Taxonomy for AI/ML](https://github.com/CycloneDX/cyclonedx-property-taxonomy/blob/main/cdx/ai-ml.md)_ | | 2.2.2.(ii) | If publicly known, list private datasets obtained from other third parties | Publicly known, third-party, private dataset information would be provided similarly to public datasets.

See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ | | 2.2.2.(iii) | General description of non-publicly known private datasets obtained from third parties | Non-publicly known, third-party, private dataset information would be provided similarly to public datasets.

See [template mappings](#template-mappings), Section 2.1 _"Publicly available data"_ (above) | _See referenced section._ |