Modcon.ai
Home
Modcon Group
About Modcon.AI
Solutions
Contact Us
Resources
Testimonial
Modcon.ai
Home
Modcon Group
About Modcon.AI
Solutions
Contact Us
Resources
Testimonial
More
  • Home
  • Modcon Group
  • About Modcon.AI
  • Solutions
  • Contact Us
  • Resources
  • Testimonial
  • Home
  • Modcon Group
  • About Modcon.AI
  • Solutions
  • Contact Us
  • Resources
  • Testimonial

Model-AI Chemometrics

Automatic and Semi-Automatic Chemometric Model Development for Laboratory and Process NIR Analyzers

 

Near-infrared spectroscopy has become an established analytical technique for laboratories and continuous industrial processes. A single NIR spectrum may contain information related to several physical and chemical properties, allowing one instrument to estimate composition, quality and process performance without the lengthy procedures associated with many conventional laboratory methods.


The spectrometer, however, is only part of the measurement system. An NIR analyzer does not normally measure a property such as density, octane number, sulphur, water, distillation point or product composition directly. It records the interaction between near-infrared radiation and the sample. A mathematical model then converts the resulting spectral information into an estimated property value.


The performance of the analyzer therefore depends on three connected elements:

  1. The quality and stability of the spectrometer
  2. The quality of the reference laboratory data
  3. The quality, relevance and maintenance of the chemometric model


The third element is often the most difficult to manage over the operating life of the analyzer. Modcon Systems Ltd. has released an updated version of Model-AI, its chemometric software for developing, validating, maintaining and automatically updating models used with laboratory and process NIR analyzers. The new release is available in both automatic and semi-automatic modes, allowing the same platform to support routine users, laboratory specialists and experienced chemometricians.


The software automates key activities such as data preparation, model training, validation and intelligent outlier detection. Its purpose is to make advanced modelling procedures accessible to operators with different levels of experience while maintaining the controls required for reliable analytical use.


Why NIR models require continuous attention

A chemometric model represents relationships found within a defined set of spectra and reference laboratory results. It is not a permanent mathematical truth.


The original calibration may perform very well during commissioning. Over time, however, the process can move beyond the conditions represented in the original dataset.


Typical reasons include:

  • changes in crude oil, feedstock or raw-material origin
  • variation in product formulation
  • seasonal changes in feed composition
  • introduction of new components or additives
  • changes in process temperature or pressure
  • catalyst ageing
  • equipment modifications
  • spectrometer maintenance or component replacement
  • changes in sample presentation
  • contamination or ageing of optical windows
  • differences between laboratory methods, instruments or operators


A model developed from a narrow dataset may produce excellent statistics during initial testing but perform poorly when the real process begins to explore a wider operating envelope. This is one of the less glamorous truths of process spectroscopy: an impressive calibration report does not guarantee a useful long-term analyzer. The reliability of an indirect analytical method depends strongly on the data used for model development and testing. Outliers, uninformative samples and unrepresentative datasets can reduce performance, while representative and influential samples can substantially improve the model. For this reason, model maintenance must be treated as part of analyzer lifecycle management rather than as an occasional software exercise.


The traditional model-maintenance problem

Conventional chemometric model development is normally performed manually by a specialist. The work may include:

  • importing spectra and laboratory results
  • matching spectra with corresponding reference samples
  • checking timestamps and sample identities
  • reviewing missing or incorrectly formatted data
  • selecting the property to be modelled
  • identifying unsuitable spectra
  • selecting preprocessing methods
  • dividing the dataset into calibration, validation and test groups
  • identifying outliers
  • selecting the number of latent variables or factors
  • comparing alternative model structures
  • evaluating statistical and analytical performance
  • exporting the approved model
  • deploying it to an analyzer
  • monitoring its performance after deployment


The procedure is technically demanding and can be slow. Manual model building typically requires repeated cycles of data selection, preprocessing, training, outlier investigation, model rebuilding and statistical review.


In many industrial facilities, the site does not maintain a full-time chemometrics specialist. Model updates may therefore depend on the analyzer supplier, a third-party consultant or a central corporate laboratory.


This arrangement can create several practical problems:

  • model updates are delayed
  • new reference samples remain unused
  • the analyzer gradually becomes less representative of the process
  • minor model problems accumulate
  • local personnel lose confidence in the result
  • the analyzer is used for monitoring but not control
  • laboratory testing remains unnecessarily high
  • ownership of model performance becomes unclear


The attached Model-AI technical material identifies the same operating problem: continuous changes in processes and raw materials require frequent model updates, while traditional updates can involve significant cost, long service procedures, third-party involvement and highly qualified chemometric personnel.


What is Model-AI?

Model-AI is a software environment developed to simplify the creation, validation, deployment and management of mathematical models for correlative analyzers.


It is designed to support the complete relationship between:

  • spectra generated by an NIR analyzer or laboratory spectrometer
  • corresponding reference laboratory results
  • chemometric model development
  • model validation
  • model export and deployment
  • subsequent model maintenance and updating


The platform is not limited to a single fixed instrument architecture. It supports model-building procedures for different types of correlative analyzers and can therefore be applied across laboratory and process-analysis environments.


The updated release introduces two distinct operating approaches:

  • Automatic mode
  • Semi-automatic mode


This is important because industrial users do not all require the same level of control.

A production laboratory may want the software to perform most routine modelling work automatically. A central analytical group may prefer to inspect data, adjust selections and compare candidate models. A senior chemometrician may require direct control of certain modelling decisions while still benefiting from automated data checks and validation.


The new architecture supports these different working practices without requiring separate software packages.


Automatic mode

Automatic mode is intended for routine model development and maintenance where the user wants the software to manage the principal modelling decisions within predefined technical and quality limits.


A typical automatic workflow includes:

  1. Data import
  2. Spectrum and laboratory-result matching
  3. Data-quality checks
  4. Property selection
  5. Screening for missing or invalid records
  6. Identification of unusual spectra or reference values
  7. Dataset partitioning
  8. Model training
  9. Cross-validation
  10. Selection of an appropriate model structure
  11. Statistical assessment
  12. Generation of the candidate model
  13. Comparison with the existing model
  14. Approval or controlled deployment


The aim is not simply to make the process faster. It is to make routine modelling more consistent. Two specialists may make different decisions when reviewing the same dataset. Automatic procedures reduce this operator-to-operator variation by applying the same screening, training and validation logic each time.


This is particularly useful where several analyzers, laboratories or production sites must be maintained according to a common method.


Automated data screening

Before model training begins, the software checks the available data.


Potential problems may include:

  • spectra without corresponding laboratory results
  • laboratory results without spectra
  • duplicate records
  • incorrect sample identifiers
  • invalid dates or timestamps
  • incorrectly formatted numerical values
  • missing property values
  • samples outside configured operating limits
  • spectra with abnormal intensity
  • saturated detector regions
  • excessive spectral noise
  • inconsistent wavelength ranges
  • incomplete spectrum files


Model-AI includes reporting functions for laboratory results that do not have corresponding spectra and for formatting errors.


These checks are not administrative decoration. A wrongly matched laboratory result can damage the model more seriously than a visibly poor spectrum because the data may appear technically valid while teaching the model an incorrect relationship.


Automatic outlier detection

Outlier handling is one of the most sensitive parts of chemometric model development. An outlier may result from:

  • an incorrect laboratory value
  • a wrongly identified sample
  • contamination
  • a process upset
  • poor sample handling
  • an instrument fault
  • a genuine but unusual process condition


Automatically deleting every unusual point would be unsafe. Some apparently abnormal samples may be the most valuable points in the dataset because they extend the model into an important operating region.

Model-AI therefore uses intelligent outlier-detection procedures to identify and manage anomalous data while supporting model accuracy and reliability.


The software can flag suspicious samples, assess their influence and determine whether they should be excluded, retained or referred for user review according to the configured operating mode.


In automatic mode, the system follows approved rules. In semi-automatic mode, the user can inspect the relevant sample and make the final decision.


Automated model training

The software trains candidate models using the accepted dataset and selected property. During training, it evaluates the relationship between spectral variation and the reference property. Depending on the configured methodology, the procedure may consider:

  • spectral preprocessing
  • wavelength selection
  • data scaling
  • calibration factors
  • latent-variable selection
  • model complexity
  • training error
  • validation error
  • sensitivity to individual samples


The objective is not to produce the lowest possible calibration error. A model can fit the training data extremely well and still perform badly on new samples.


A useful industrial model must balance:

  • accuracy
  • robustness
  • simplicity
  • sensitivity
  • stability
  • transferability
  • resistance to overfitting


Model-AI therefore includes separate stages for data screening, training-data division and initial model training.


Automatic validation

Validation checks whether the model can predict samples that were not used directly to fit it. The software can divide available data into groups for:

  • model training
  • model validation
  • independent testing


Chronological division may be particularly useful for process data. Random division can place nearly identical consecutive process samples into both training and validation sets, creating an unrealistically favourable result.


A chronological test gives a more demanding indication of how the model may perform when applied to later production data.


Typical model-performance indicators may include:

  • coefficient of determination
  • calibration error
  • validation error
  • test-set error
  • bias
  • slope
  • residual distribution
  • maximum individual error
  • repeatability
  • number of model factors
  • proportion of samples excluded
  • stability across the property range


No single statistic should be treated as sufficient. A high coefficient of determination may look attractive, but it does not by itself prove that the model is suitable for process control. The result must also be assessed against the reference-method uncertainty, required process accuracy and intended operational use.


Semi-automatic mode

Semi-automatic mode retains the structured Model-AI workflow but allows the user to intervene at defined stages.


This mode is intended for users who require more visibility and control, including:

  • chemometricians
  • analytical chemists
  • process-analysis specialists
  • central laboratory teams
  • analyzer engineers
  • research and development groups


The user can review or adjust decisions related to:

  • sample selection
  • property selection
  • outlier acceptance or rejection
  • calibration and validation datasets
  • chronological or alternative data splitting
  • preprocessing choices
  • spectral regions
  • model factors
  • retraining
  • candidate-model comparison
  • final model approval


The purpose of semi-automatic operation is not to return to a fully manual procedure. Routine data checking, calculation and statistical analysis remain automated. The user concentrates on decisions where process knowledge or analytical judgement adds value.


This is often the most appropriate mode for difficult industrial applications. For example, a new crude oil may appear as a spectral outlier. A purely statistical decision might reject it. A refinery specialist may know that the crude will become a regular feed component and should therefore remain in the calibration dataset. The combination of automated analysis and informed human review is more useful than either one in isolation.


Automatic model updating

The most significant capability of the updated Model-AI release is support for automatic model updating.

A process NIR model should evolve as the process evolves. This does not mean continuously replacing the model whenever a new laboratory result becomes available. Uncontrolled adaptation could allow poor laboratory data, temporary disturbances or instrument faults to enter the model.

Automatic updating must therefore be governed by a controlled sequence.


A controlled update cycle

A robust automatic model-update cycle can include the following stages.


1. Collection of new spectra

The laboratory or process NIR analyzer continues to collect spectra during routine operation.

For process applications, each spectrum should be associated with relevant metadata such as:

  • analyzer identification
  • timestamp
  • sample or stream identification
  • operating mode
  • product grade
  • process temperature
  • process pressure
  • flow condition
  • maintenance status
  • alarm status


This contextual information helps prevent unsuitable operating periods from entering the model-maintenance dataset.


2. Association with reference results

Where a physical sample is collected for laboratory testing, the software associates the reference result with the corresponding spectrum or representative spectral period. This stage requires careful time alignment.


The timestamp of laboratory analysis is rarely the correct timestamp for model matching. The relevant time is normally when the sample was taken from the process. Sample-line transport delay, analyzer response time and process residence time may also need to be considered. Incorrect time alignment is a common source of model degradation in process spectroscopy.


3. Data-quality verification

Before a new sample is accepted, Model-AI checks whether:

  • the spectrum is complete
  • the analyzer was operating normally
  • the reference result is available
  • the property value is plausible
  • the sample is correctly matched
  • the data format is valid
  • the result lies within the configured application range
  • the sample contributes useful new information

A new result should not enter the model merely because it exists.


4. Applicability assessment

The software determines whether the new spectrum is adequately represented by the current model.

The sample may fall into one of several categories:

  • normal sample within the existing calibration space
  • useful sample that improves population density
  • boundary sample that extends the model range
  • influential sample representing a new process condition
  • suspicious outlier requiring investigation
  • invalid sample unsuitable for modelling

This distinction is important. Hundreds of nearly identical samples may add little value, while a small number of well-selected boundary samples can materially improve model coverage.


5. Candidate model generation

When sufficient valid new data are available or a configured update condition is reached, the software generates a candidate model.

The candidate model is developed using the existing approved data together with selected new samples. The process should avoid replacing the historical dataset entirely, as this may cause the model to forget earlier but still valid operating conditions.

The update procedure must balance:

  • adaptation to current process conditions
  • preservation of historical process coverage
  • avoidance of unnecessary model complexity
  • protection against poor-quality new data


6. Comparison against the approved model

The candidate model is then compared with the model currently in operation.

The comparison may consider:

  • performance on the original validation dataset
  • performance on recent samples
  • independent test-set error
  • bias
  • slope
  • maximum prediction error
  • number of factors
  • number and identity of excluded samples
  • stability across different process regimes
  • performance near specification limits

A candidate model should not be approved merely because it performs better on the newest samples. It must not produce unacceptable deterioration elsewhere.


7. Approval and deployment

Depending on the configured governance level, the candidate model may be:

  • deployed automatically
  • held for operator approval
  • held for specialist approval
  • rejected automatically
  • returned for further investigation

This is where the two operating modes are particularly useful.

In automatic mode, deployment can proceed when all predefined acceptance criteria are satisfied.

In semi-automatic mode, the system presents its findings and recommendations to an authorised user before deployment.


8. Post-deployment monitoring

After installation, the new model remains under observation.

The system should monitor:

  • prediction bias
  • laboratory comparison
  • residual trends
  • frequency of out-of-model samples
  • prediction stability
  • process consistency
  • analyzer diagnostics

If performance deteriorates, the model can be rolled back to the previously approved version.

Model versioning and traceability are therefore essential parts of automatic updating.


Updating should be triggered by evidence

Automatic model updating does not necessarily mean updating according to a fixed calendar. A model should be updated when the data show that an update is justified.


Possible triggers include:

  • accumulation of a specified number of validated new samples
  • persistent prediction bias
  • increased prediction residuals
  • a growing number of out-of-model spectra
  • introduction of a new product grade
  • change of feedstock
  • change in formulation
  • major analyzer maintenance
  • optical-component replacement
  • process modification
  • laboratory-method change
  • deterioration beyond an approved performance threshold


Different applications require different update policies. A blending analyzer working with frequently changing feed components may require regular updates. A stable utility-stream application may operate successfully for a long period without model modification. Automatic updating should therefore be configurable, evidence-based and application-specific.


Laboratory NIR applications

In a laboratory environment, Model-AI can support model development using spectra acquired from benchtop or at-line NIR instruments.


Typical laboratory applications may include:

  • raw-material identification
  • product-quality screening
  • formulation analysis
  • incoming-material inspection
  • fuel-property estimation
  • chemical-composition analysis
  • research and development
  • transfer of conventional laboratory methods to faster NIR methods


The laboratory workflow usually provides better control over:

  • sample temperature
  • sample presentation
  • optical path length
  • mixing
  • measurement repetition
  • reference-method timing


This can simplify model development, but it does not remove the need for careful data management. Laboratory datasets may still contain hidden differences caused by:

  • different operators
  • different sample cells
  • inconsistent sample preparation
  • ageing samples
  • temperature variation
  • different reference instruments
  • changed laboratory methods
  • transcription errors


Model-AI provides a common structure for collecting, screening, modelling and validating these data. Automatic mode can be used for routine models and well-defined applications. Semi-automatic mode is more appropriate during method development or where the analyst needs to investigate specific spectral and statistical effects.


Process NIR applications

Process NIR analysis presents additional challenges because the analyzer operates continuously under changing plant conditions.


The optical system may be exposed to:

  • changing temperature
  • changing pressure
  • variable flow
  • multiphase samples
  • bubbles
  • suspended solids
  • coating or fouling
  • vibration
  • changing background composition
  • rapid feedstock transitions


The model must distinguish useful compositional information from unrelated physical variation.

For this reason, a process model should be developed from data that represent the real installation, not only ideal laboratory conditions.


The architecture presented in the supplied Model-AI material links the process, laboratory, spectra, software, model and process analyzer in a continuous analytical chain.


This architecture reflects an important principle: the process analyzer, laboratory and modelling system cannot be managed independently.


The laboratory provides reference values. The analyzer provides spectra. The process provides operating variability. Model-AI combines these elements into a maintained analytical method.


Model transfer from laboratory to process

A laboratory model may provide a useful starting point for a process application, but direct transfer is not always straightforward.


Differences may exist in:

  • instrument optical design
  • detector characteristics
  • wavelength resolution
  • optical path length
  • sample temperature
  • sampling geometry
  • sample homogeneity
  • instrument environment
  • spectral preprocessing
  • wavelength calibration


A laboratory spectrum collected from a carefully prepared sample may not be identical to the spectrum measured continuously in a process cell. Model transfer may therefore require:

  • standardisation samples
  • bias correction
  • slope correction
  • instrument standardisation
  • additional process samples
  • recalibration
  • wavelength alignment
  • separate process validation


Model-AI can assist by maintaining the relationship between the laboratory reference data and the spectra from the actual analyzer on which the model will be used.


Data diversity matters more than data volume

Large datasets are useful only when they contain relevant variation. A calibration containing several thousand nearly identical samples may be less robust than a smaller dataset covering:

  • different feedstocks
  • different grades
  • normal operating variation
  • process transitions
  • low and high property values
  • seasonal conditions
  • expected disturbances
  • specification boundaries


The updated software is designed to handle large datasets and complex spectra, but scalability should not be confused with indiscriminate data accumulation.


Automatic data selection should favour informative samples rather than simply accepting every available record.


This also reduces unnecessary model complexity and avoids overweighting the most common operating condition.


Reference-method uncertainty

A chemometric model cannot be more reliable than the reference data used to train it. Every laboratory method has uncertainty arising from:

  • repeatability
  • reproducibility
  • sample preparation
  • instrument calibration
  • operator technique
  • environmental conditions
  • method limitations


Where the laboratory method itself has substantial variation, the NIR model may appear to have a prediction error even when part of the difference is caused by laboratory uncertainty.


Model performance must therefore be evaluated in the context of the reference method. Before model development, it is advisable to establish:

  • method repeatability
  • expected inter-laboratory variation
  • sample stability
  • correct sample-handling procedure
  • appropriate number of repeat measurements
  • procedure for resolving questionable results


Automatic modelling can process data efficiently, but it cannot repair an unreliable reference method by statistical enthusiasm.


Model governance and traceability

Where NIR results are used for process control, product release, blending or commercial decisions, model governance becomes essential. Each approved model should have a traceable record containing:

  • model identifier
  • version number
  • analyzer identifier
  • property name
  • measurement range
  • dataset version
  • sample list
  • preprocessing configuration
  • modelling method
  • validation results
  • excluded samples
  • reason for exclusion
  • approval status
  • approving user
  • deployment date
  • previous model version
  • rollback status


Automatic updates should create a new controlled version rather than silently modifying the active model. The software should also distinguish between:

  • development models
  • candidate models
  • approved models
  • deployed models
  • retired models


This prevents an experimental calibration from being mistaken for the production model.


Human oversight remains important

Automation reduces dependence on specialist resources, but it does not eliminate the need for process and analytical knowledge. A model can pass statistical tests while remaining operationally unsuitable. 


For example:

  • the dataset may not cover a critical product grade
  • the laboratory method may have changed
  • apparent model drift may actually be analyzer fouling
  • a new cluster may represent a genuine feedstock change
  • the model may be accurate overall but weak near a specification limit
  • a low validation error may result from excessive similarity between training and test samples


The purpose of Model-AI is therefore not to remove people from the decision process. It is to automate repetitive work, identify issues early and present the relevant information in a structured form.

Automatic mode supports routine, governed operation. Semi-automatic mode allows informed intervention when the application requires judgement.


Technical platform

The supplied product specification indicates support for Windows 10 or later and selected Linux environments, including Astra Linux and Ubuntu. The software uses an SQLite database and supports CSV and SPC spectrum-file formats, with AMO model output.


This allows Model-AI to be used in different laboratory and industrial-computing environments while supporting common spectral-data exchange formats.


The software is intended to sit within the wider analyzer-management architecture rather than operate as an isolated statistical package.


From analyzer installation to analytical lifecycle management

Historically, many NIR projects concentrated on instrument selection and initial calibration. Once the analyzer was commissioned, model maintenance was treated as a secondary activity.


This approach is no longer adequate for modern process applications. A process NIR system should be managed as a lifecycle:

  1. Define the analytical objective
  2. Establish the reference method
  3. Collect representative samples
  4. Build the initial model
  5. Validate under real conditions
  6. Deploy with controlled versioning
  7. Monitor model performance
  8. Collect new representative data
  9. Update when justified
  10. Revalidate and redeploy
  11. Maintain complete traceability


Model-AI provides the software framework for this cycle. The addition of automatic and semi-automatic modes allows the system to serve both routine industrial users and specialists. Automatic updating extends this approach further by allowing the model to adapt to process changes while remaining subject to defined quality and approval rules.


Conclusion

NIR spectroscopy offers fast, non-destructive and multi-property analysis, but its long-term value depends on the quality and maintenance of its chemometric models.


Changing feedstocks, process conditions, product grades and analyzer environments mean that a model cannot simply be built once and forgotten. It must be monitored, challenged and updated using reliable laboratory data and representative process spectra.


The updated version of Model-AI addresses this requirement through:

  • automatic model development
  • semi-automatic expert-guided modelling
  • structured data screening
  • intelligent outlier detection
  • automated training and validation
  • controlled model comparison
  • automatic model updating
  • model versioning and deployment control
  • support for both laboratory and process NIR analyzers


The result is a more practical approach to chemometric model lifecycle management. Automatic mode makes routine model maintenance faster and more consistent. Semi-automatic mode preserves expert control where process knowledge and analytical judgement are required. Together, they allow users to maintain reliable NIR models without turning every update into a small research project.

In industrial analytics, the analyzer produces the spectrum. The laboratory provides the reference. The model turns both into a useful measurement. Keeping that model current is not optional; it is part of the measurement itself.

  • Home
  • Modcon Group
  • CDU Optimization Suite
  • AI Energy Conservation
  • Process Health Analysis
  • AI Medical Technology
  • Hydrogen Production
  • Model-AI Chemometrics

10 Orange St., Haymarket London WC2H 7DQ UK

+44-204-5771737

www.modcon-systems.com

Copyright © 2026 Modcon Systems Ltd. - All Rights Reserved.

www.modcon-systems.com

This website uses cookies.

We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.

Accept