Skip to main content
Research Guidelines

American Meat Science Association Research Guidelines for Cookery, Sensory Evaluation, and Instrumental Texture Measurements of Meat

Authors
  • Rhonda K. Miller orcid logo (Texas A&M University)
  • Linda S. Papadopoulos (LP & Associates, LLC)
  • M. Wes Schilling (Mississippi State University)
  • Tommy L. Wheeler orcid logo (USDA, Agricultural Research Service)
  • Rachel A. Gray (Walmart)
  • Michael E. Dikeman (Kansas State University)
  • James R. Claus (University of Wisconsin–Madison)
  • Casey M. Owens (University of Arkansas)
  • D. Andy King (USDA, Agricultural Research Service)
  • Keith E. Belk orcid logo (Colorado State University)
  • Phillip Bass (University of Idaho)
  • Chris Calkins orcid logo (University of Nebraska–Lincoln)
  • Mark Miller (Texas Tech University)
  • Steven D. Shackelford (USDA, Agricultural Research Service)
  • Bridget Wasser (Wasser Consulting Services, LLC)

Abstract

The American Meat Science Association (AMSA) historically has published recommendations for procedures to harmonize methodology when conducting research related to sensory traits of meat. Consistency in research methodology has provided greater understanding of sources of variation in meat-sensory traits and facilitated comparisons among studies in the literature. The guideline s have expanded significantly since the AMSA (1978) document was published. The guidelines are not a standard and leave the final decisions on methodology to researchers so that they can best address their research objectives. Recommendations require periodic updating. These guidelines were expanded to cover additional species, processed meats, discriminative testing methods, and new developments in cookery methodology. The guidelines provide methodology for meat cookery; sensory environment, product and panelist controls; discriminative, descriptive, and consumer sensory methods; and methods for measuring meat texture for whole-muscle and processed meat products.

Keywords: guidelines, sensory, cookery methods, meat, discriminative sensory, descriptive sensory, consumer sensory, texture measurements

How to Cite:

Miller, R. K., Papadopoulos, L. S., Schilling, M. W., Wheeler, T. L., Gray, R. A., Dikeman, M. E., Claus, J. R., Owens, C. M., King, D., Belk, K. E., Bass, P., Calkins, C., Miller, M., Shackelford, S. D. & Wasser, B., (2026) “American Meat Science Association Research Guidelines for Cookery, Sensory Evaluation, and Instrumental Texture Measurements of Meat”, Meat and Muscle Biology 10(1): 20205, 1-104. doi: https://doi.org/10.22175/mmb.20205

Rights:

© 2026 Miller, et al. This is an open access article distributed under the CC BY license.

Funding

Name
American Meat Science Association Development Council

478 Views

90 Downloads

1 Citations

Published on
2026-06-20

Peer Reviewed

Introduction

In 1978, the American Meat Science Association (AMSA) first published Guidelines for Cookery and Sensory Evaluation of Meat (AMSA, 1978). During the next 17 y, these guidelines were useful to AMSA members and nonmembers involved in meat cookery and sensory evaluation. In 1995, an update was published by the AMSA and the National Live Stock and Meat Board (AMSA, 1995). A 12-member AMSA committee expanded and updated the guidelines to provide greater consistency, accuracy, and relevancy to cookery, texture, and sensory research. Numerous changes in cooking equipment, sensory methods, and meat products have occurred since the 1995 guidelines were published (AMSA, 1995). As a result of diet/health concerns and changing consumer preferences, meat products have become leaner and there is greater variation in types of products, especially for value-added or processed meats. Additionally, meat scientists have more extensively utilized consumer sensory methods as a research tool. The cookery methods included in this update are accepted and recommended by the authors. These are not consumer cookery guidelines but research guidelines, and recommended research methods may not always produce the highest level of consumer acceptance or always emulate consumer cooking and preparation procedures. These methods are designed to control unwanted variability, to determine the most accurate answer to the questions being addressed using relevant methods, and to allow for valid comparative interpretation of research. Cookery and product-handling suggestions are included for fresh and processed meat as defined by the AMSA Meat Lexicon (Seman et al., 2018). Therefore, meat is a term used to describe protein products across species, including beef, veal, pork, lamb, goat, chicken, turkey, duck, goose, and other mammalian, avian, reptilian, amphibian, and aquatic species. Processed meats are a term used to describe meat that has been salted, cured, smoked, or fermented to enhance flavor and/or extend shelf-life. The guidelines presented here do not imply that other suitable procedures/equipment are not available but do provide procedures that are reliable and appropriate. This guideline includes a list of factors to consider when designing and planning sensory and instrumental texture/tenderness research, recommendations to consider when establishing and implementing best practices to maximize test sensitivity, as well as suggestions on testing methodologies and statistical methodologies.

Research Planning

Before initiating an experiment, the cooking and handling procedures, sensory method, instrumental texture/tenderness method, and testing parameters should be determined. Factors to consider in method selection include the following:

  • What is your hypothesis or what are you trying to learn?

  • What questions are you trying to answer (test objectives)?

  • How will the results be used?

  • How large of a difference are you trying to detect?

  • How much variability is there within and between samples?

  • How many factors do you need to test to adequately address the test hypothesis?

Regardless of what products are being tested or what test method used, successful testing is based on identifying and controlling/accounting for as many sources of variability as is reasonable based on the objectives of the test. There will be situations where variability is naturally large due to processing parameters or the nature of the product (e.g., fermented sausages). In these cases, it is critical to account for, minimize, and accommodate variability through sound experimental design. Additionally, targeting the right respondent in the sensory test is critical for getting accurate responses and for application of results.

The diagram shown in Figure 1 may be useful in selecting the most appropriate sensory testing method/s. If, based on preliminary work, the sensory differences among treatments are not expected to be detectable, the lack of significant differences can be verified using discrimination or descriptive analysis methods. If, however, the sensory differences are expected to be detectable, consumer testing methods would be more appropriate. It is important to determine if the differences are detectable to consumers, and if detectable, how they affect consumer acceptability.

Figure 1.
Figure 1.

Flowchart for determining the appropriate test method(s) for sensory testing.

Quantitative sensory methods can be grouped into 3 primary categories: (1) discrimination, (2) descriptive analysis, and (3) affective (consumer) (Meilgaard et al., 2025). Discrimination methods can include either trained panelists or untrained consumer panelists, depending on the test objectives. If the test objectives are to determine with a high degree of certainty if treatment differences are significant, trained panelists are suggested. Trained panelists are carefully selected, highly trained, and hypercritical as compared to average consumers. When consumers are the panelists in discrimination tests, consumers are not as critical and may not detect differences. The selection of a testing method should be based on the objectives of the study. Data should be interpreted based on the sample population that is used for the study.

Descriptive methods use trained panelists. The amount of training and experience of panelists impacts their ability to discriminate (Chambers et al., 2004; ASTM MNL90-2ND-EB, 2023). As the amount of training and experience increases, panelists can detect smaller differences in attributes between samples. The amount of training should be reported, and data should be presented based on the panelists’ level of training. Descriptive tests are used to quantify the level of an attribute within the samples. Depending on the testing method selected, scales and attributes may vary, but each test is used to determine if samples differ in sensory attributes.

Consumer evaluation can be either qualitative or quantitative. In qualitative consumer methods, consumers provide opinions and perceptions of products (Meilgaard et al., 2025). These data are very valuable, but they do not give quantitative results that can be statistically analyzed. Consumer quantitative testing provides an avenue to measure consumer opinions using questions on a ballot with a scale that is or can be converted to numerical values for statistical analyses. This provides a method of quantifying consumers’ responses and determining differences. There are 2 commonly used consumer quantitative testing methods, which are central-location tests (CLT) and home-use tests (HUT). CLT are conducted in highly controlled environments and are not normal conditions for food consumption (Meilgaard et al., 2025). HUT provide the opportunity to include in-home preparation, family opinions, and environment for normal consumption of the product. When selecting the test to use, sensory professionals should always remember that consumer panels differ from trained panels. Although members of a trained panel are consumers, their opinions and preferences may not be representative of the general population. Trained panelists should, therefore, never be asked to respond with their opinions on preference or liking (Meilgaard et al., 2025). More information on the testing methods and considerations is provided in the appropriate sections. Meat scientists are strongly encouraged to understand each testing method and the strengths and weaknesses of each method when determining the method that provides the best procedures to meet the test hypothesis and objectives.

Mechanical measurements of texture/tenderness provide quantitative assessments of meat products. These methods vary across species and processed meat products. While these methods do not provide a human response to the product, they provide a standardized measurement of the product structure, defined as rheological properties. Relationships between mechanical measurements and sensory assessment have been examined. Commonly used methods will be discussed, and the strengths and weaknesses of each method will be provided.

Sample Collection/Preparation

Postmortem processing procedures, including rate of chilling, aging (duration and conditions), fabrication, and freezing, have been shown to affect sensory properties of whole-muscle and processed meats. Processed meats can also be impacted by countless other variables, such as ingredient supplier, ingredient quality, product processing conditions, and storage time and temperature of the finished product. The guidelines provide recommendations for addressing such aspects of research, but it is realized that the overall objectives of projects may dictate variation in these factors. It is recommended that postmortem process procedures be standardized and documented for studies. When results are either across institutions or time, it is important to ensure uniformity during and across study subsets.

Typically, steaks, chops, filets, and roasts should not be sized (cut to a specific dimensions or weight) unless preslaughter treatment or postmortem processing affects the intended size. There is merit in “sizing” meat cuts if variation in cooking procedures is the major focus of the project. Examples of “sizing” meat cuts include removal or standardization of bone content, connective tissue, subcutaneous fat, and standardization of thickness, weight, and shape. Ground beef patties should have very close controls on weight and thickness. Variation in patty manufacturing parameters should be considered and standardized, including grinding size, mixing time, batch size within equipment, temperature during grinding, mixing and patty formation, and cooling or freezing conditions. Like ground beef patties, manufacturing parameters for processed meats such as frankfurters, sausage links, ham, bacon, deli meat, and dry-cured meats should be standardized and closely monitored to minimize the introduction of unintended variance. Specific variables to consider include temperature of meat at manufacture or purchase, postmortem age of raw material, ingredient source and production lot, line-to-line and batch-to-batch variation, time of production (beginning, middle, end of run), slicing consistency, line speed, and others that may affect the consistency of products. Before starting research, consider all possible sources of variance and take the necessary steps to control and/or standardize those sources so that resulting data are relevant and accurate. Scientists should consider using a minimum of 3 replications to account for day, production run, or processing effects that can then be accounted for as random or fixed effects in statistical models. However, statistical power tests should be conducted to validate the number of replicates needed, and sample size should be large enough that statistical significance is found with both biological and practical meaning.

Selection of samples

Researchers are strongly advised to consult with statisticians and/or perform appropriate statistical power tests for determination of experimental sample size and number of experimental units. A number of parameters, including an estimate of variability in texture (instrumental or sensory), flavor, or other traits of interest are needed to determine appropriate sample size. Many statistics textbooks provide guidelines for estimating an appropriate sample size to control Type II error. Likewise, there are many online options. It is more fiscally and scientifically sound to have adequate sample size than to use marginal numbers of observations and not detect differences that might truly exist (not controlling for Type II error). As a rule, regardless of the product or species being tested, all samples within a study should be representative of the products or processes under study. Products should be produced under the same conditions if possible. Proper sampling within muscles is of critical importance. Steaks or chops within a muscle that are to be assigned to different treatments should be randomized (if there is no known location variation) or blocked (if there is known location variation) to control for variance. Location can be accounted for either as random effects or fixed effects, depending on the design. If each carcass or cut represents 1 replicate of 1 treatment, steak location within a muscle should be standardized. For example, the first steak from the rib end of the short loin could be assigned to sensory analysis, and the second steak should be assigned to shear force evaluation tests. Some noncomminuted processed meats may need to be treated more like fresh meats when sourcing samples. For example, bacon would be randomized by belly and slice origin within the belly (a slice from the shoulder portion of the belly has different characteristics than a slice from the center or flank). A comminuted processed meat product, such as a sausage patty, would be randomized by and/or within production run/s.

Most research that has the goal of determining production, antemortem, or postmortem treatment effects on sensory traits utilizes a small number of muscles. Greater emphasis, however, has recently been placed on characterizing and marketing a greater number of muscles; therefore, it may be appropriate to evaluate treatment effects on multiple muscles. Location effects within muscle can contribute significantly to source of variation in some muscles. In addition, muscle selection should be based on experimental questions for the study. Data indicate that the relationships in tenderness among muscles in the same animal vary from relatively low to moderately high; therefore, treatment effects in one muscle may or may not be representative of similar treatment effects in other muscles (Ramsbottom et al., 1945; McKeith et al., 1985; Belew et al., 2003).

Processing conditions need to be taken into consideration when selecting processed meat samples. Multiple controls may be included if they are produced in the same plant and on the same line and equipment as other products in the study. Production should be in a steady state before any product is sampled for testing, and samples from a production shift should be pulled at 3 different times to ensure that the products are representative of the shift. The number of production dates or lots obtained will depend on the objective of the research. Standardization of the number of production lots and sample age are needed to accurately assess differences from product to product and should be representative of the product for which the results will be inferred.

Time postmortem for processing into cuts

Cuts should not be removed before rigor mortis is complete unless development of rigor mortis is an objective. The timing of rigor mortis completion may vary among species and with chilling conditions. Beef and lamb carcasses should not be ribbed and fabricated into cuts before 24 h postmortem, and pork carcasses should not be fabricated for approximately 18 h. Poultry and fish should follow recommendations based on completion of rigor. Timing is dependent on the research objectives and schedules of commercial facilities if the product is obtained commercially. For additional postmortem aging time, cuts should be vacuum packaged (unless dry aging is a component of research objectives). If a study includes alternative packaging systems in the experimental design, meat should be packaged to represent the objectives of the experiment. Extensive research has been conducted on the effects of aging on sensory properties of meat and meat products. Researchers are strongly advised to understand the effects of aging within their product prior to determining aging time/s for a study. The following postmortem times are recommended across species unless aging time is a factor of the study:

  • Beef: between 10 and 28 d, with 14 d being optimum for most research.

  • Pork: between 5 and 14 d, with 10 d being optimum for most research.

  • Lamb: between 7 and 14 d, with 7 d being optimum for most research.

  • Poultry: poultry meat deboning times vary. Two hours postmortem is common due to industry practices, 4 to 6 h postmortem or longer is recommended to detect differences between treatments that do not involve processing treatments, and aging poultry meat for 24 h postmortem prior to cooking or freezing is recommended.

  • Fish: due to variation in processing and slaughtering practices for commercially raised or wild-caught fish species, conditions including initial storage, freezing, and thawing, length of time for each respective stage, temperatures, and packaging should be defined and emulate commercial practices. Conditions should be standardized within a study.

  • Processed meats: 14 d postprocessing to allow for flavor equilibration.

These aging times are optimum for detecting differences among breeds and management systems, as well as antemortem and postmortem treatments without aging time functioning as a confounding factor, but they should not preclude research with unique postmortem treatments. Most aging effects on tenderness and flavor equilibration are complete by the above recommended times, but the industry average is considerably longer with wide variation within and among retail and foodservice products.

Meat dimension variables

Recommended thicknesses of steaks, chops, filets, patties, and processed meats are as follows: beef steaks (dry heat): 2.54 cm; beef steaks (moist heat): 1.9 to 2.54 cm; lamb chops (dry heat): 2.54 cm; and pork chops (dry heat): 2.54 cm. Whole poultry filets should be at least 2.54 cm in thickness or greater to represent the natural thickness of the filet. Due to portioning filets, samples may be thinner if product is marketed sliced such as thin-sliced chicken breasts, and a standard thickness should be defined.

Fish filet thickness depends on species, and for larger species, thickness should be defined based on common use of the product. Beef patties (also recommended for pork and poultry) should be between 0.95 and 1.10-cm thick for 91.5-g patties and 1.10 to 1.27 cm for 113.5-g patties. For specialty beef patties, thickness should be defined by product type such as smashed patties or hand-formed patties, and processed meats should be representative of how they would be utilized for foodservice or consumer preparation.

These are recommendations, and thinner or thicker products can be used as long as thickness is standardized within the study. If products are too thin, it is difficult to accurately measure the internal temperature during cooking. Cut thickness influences heat transfer, protein and lipid heat denaturation, and moisture loss, all of which have subsequent effects on sensory properties.

A cutting guide should be used so that meat products are uniform in thickness. Commercial meat slicers also can be used to generate fresh product from carcasses or cuts. When the product must be frozen, uniform thickness can be obtained by freezing and cutting with a saw or slicer.

Because it is impractical to list all the various roast cuts from all species, the latest editions of the Meat Buyer’s Guide (Meat Institute, 2025) and the Institutional Meat Purchase Specifications (USDA, 2014) provide common sizes of products merchandized in retail and foodservice outlets. Before cooking, retail cuts should be trimmed of excess external fat. The following are minimum weights and thicknesses for roasts: beef:1.5 kg, 5.0-cm thick; pork: 1.0 kg, 5.0-cm thick; lamb: 0.5 kg, 5.0-cm thick; and poultry: either whole, halved, or quartered carcasses as defined in the experimental design.

Freezing and frozen storage

For instrumental and/or sensory testing, steaks, chops, filets, or processed meats should be cooked, sheared, or evaluated for texture without freezing whenever possible. Freezing and thawing has been shown to decrease shear force compared to the fresh/never-frozen values for shear force due to muscle-fiber damage and ice crystal formation (Wheeler et al., 1990; Grayson et al., 2014; Beyer et al., 2024), but this can be minimalized with rapid, quick freezing and thawing at refrigerated temperature (0–4°C). When the experimental design is such that sample size or treatment protocol requires freezing of samples (multiple days of aging for example), it is acceptable as long as freezing occurs under standardized conditions across the experiment and the potential impact of freezing on results is acknowledged. As freezing samples at different aging times results in unequal frozen storage durations across treatments, these effects can be minimized by standardizing freezing and storage conditions and by randomizing aging days across sensory testing days so that each aging time is represented within a testing session.

Freezing should occur so that all samples freeze in the same amount of time by spreading samples out in one layer rather than placing them in a box and then placing them into the freezer. Likewise, rapid freezing at very cold temperatures, preferably lower than −20°C, to limit damage from ice crystal formation is recommended (Hiner et al., 1945; Qian et al., 2022). Pork and ground beef patties should be evaluated within 3 mo of frozen storage; beef and lamb steaks, roasts, and chops should be evaluated within 6 mo of frozen storage; and poultry and fish should be evaluated within 3 mo of frozen storage (AMSA, 2016). The purpose of recommended frozen storage times is to minimize lipid oxidation. Additionally, frozen sensory samples should be vacuum packaged with a high oxygen barrier film to assist in delaying lipid oxidation effects during storage.

The same freezing process recommendations apply to processed meats. The length of time that processed meats should be stored in frozen conditions depends on the characteristics of the individual product and the objective/s of the research. For example, if the researcher seeks to understand the impact of a new antioxidant, the test should be extended through the entire shelf-life of the product. In contrast, if a new flavor of sauced chicken wings is being developed, it is generally recommended to not test after 50% of the shelf-life of the product has occurred or when recognized changes in flavor and/or texture are detected. If the objective is to determine shelf-life, then an appropriate longer time would be used. For example, if the shelf-life of a fresh sausage product is approximately 17 d, you want to evaluate the product for a minimum of 17 d. Microbial level and potential pathogens should be assessed prior to sensory testing for food safety reasons.

Pre- and Postcooking Procedures

Physical state and packaging of the meat product prior to cooking

Fresh meats, steaks, chops, or filets should be vacuum packaged in oxygen-impermeable film to avoid desiccation and oxidation. Internal temperature of the meat at the initiation of cooking can affect tenderness (Berry and Leddy, 1990; Wheeler et al., 1997b). Likewise, if the temperature is lower than 2°C, the cut is prone to surface charring. The internal temperature should be recorded just prior to placing meat in the cooking environment. It is acceptable to cook ground beef patties from fresh, frozen, or thawed, depending on the objectives of the experiment as long as conditions are standardized.

Thawing conditions should be standardized for temperature, packaging, and length of time. Thawing should never be at room temperature or in water where water temperatures are above 5°C. It is preferred that raw or fresh meat products be vacuum packaged during thawing to reduce the effects of lipid oxidation. Products should be thawed in a refrigerated room with a temperature of 2 to 4°C and spaced to not overlap. The length of time for thawing should be standardized such that the internal temperature of the product reaches 2 to 4°C. Preliminary tests should be conducted to determine length of time, usually 24 or up to 48 to 72 h for larger samples. For example, vacuum-packaged beef, pork, and lamb or turkey roasts should be thawed at 2 to 4°C for 48 to 72 h or the length of time required to reach an internal temperature of 2 to 4°C. Frozen, thawed, and cooked weights should be collected for the determination of losses at different stages from the frozen to the cooked state and may include purge loss after thawing or during storage.

For processed meats, products should be treated as typically defined by the manufacturer and the objectives of the experiment for packaging, storage, and preparation (thawing). The decision to cook from the refrigerated, frozen or thawed state should be based on the product and experimental objectives. Regardless of state, temperature should be standardized at the initiation of the cooking process.

Typical storage practices for various processed meat products are presented in Table 1. Typically, products should be cooked from the intended storage state. For example, retail fried chicken products, such as chicken nuggets, are typically stored frozen in a retail plastic bag and cooked from the frozen state. Frankfurters are typically stored in a vacuum-packaged bag in refrigerated conditions, then cooked from the refrigerated state. On the other hand, premarinated individually quick-frozen poultry products may be thawed prior to cooking, and shelf-stable dried sausages can be stored and served at room temperature. Frozen, thawed, and/or cooked weights should be collected for the determination of losses at different stages from the frozen to the cooked state (if applicable).

Table 1.

Typical storage practices for processed meat products (USDA, Food Safety and Inspection Service)

Product Storage Packaging Cooking State
Breaded chicken Store as intended to be sold (either frozen or refrigerated) Plastic bag As sold
Frankfurters Refrigerated Vacuum bag Refrigerated
Premarinated, frozen (raw and cooked) meat and poultry Frozen Plastic bag Refrigerated/frozen
Premarinated (raw or cooked) meat and poultry Refrigerated Retail Over-wrap or Vacuum bag Refrigerated
Lunchmeat Refrigerated Plastic tub or vacuum bag Refrigerated
Dried sausage (i.e., pepperoni) Refrigerated Representative of manufacture storage As sold

Methods for monitoring temperature changes

Throughout the cooking process, it is imperative to monitor product and equipment temperature to prevent under, over, or uneven cooking. Both the initial and endpoint product temperatures, cooking time, equipment temperature, and product-evaluation temperature should be recorded to maintain and document consistency across samples. Temperature monitoring can be conducted using a thermocouple to determine internal meat temperature. Infrared guns or metal grill temperature monitors can be used to monitor the surface temperatures of equipment or meat exterior surfaces. Prior to use, either device should be calibrated to 0°C using ice water and to water at 100°C or appropriate boiling temperature based on altitude. Thermocouple use recommendations are presented in Figure 2. When using infrared guns, distance from the object, angle held, and environment may impact readings. The gun should be held at the manufacturer’s recommended distance from the surface at a 90° angle. Steam emanating from the surfaces may influence readings, and metal oven temperature monitors may be more accurate.

Figure 2.
Figure 2.
Figure 2.

Procedures for measuring internal meat temperature.

Evaluating degree of doneness

Because of the concern regarding food safety and for consistency in cooking that provides the greatest likelihood of detecting treatment differences, all meat in an experiment should be cooked to the same temperature, unless the effects of endpoint temperature are part of the experiment. Researchers should collect information regarding cooking times, peak/endpoint temperature, and cooking yields. Because of differences in cooking rate, the method of measuring “cooked temperature” is important. Some researchers prefer cooking to an endpoint temperature and then removing the samples from the heat, whereas others prefer to remove samples from the heat before the target endpoint is reached and then monitor temperature rise and record the peak temperature (usually for cooking methods with a high cooking rate or for heavier weight cuts). It is recommended, if cooking at a lower rate such that the postcooking temperature rise is less than 2°C, meat should be removed from the heat at the endpoint temperature. If cooking at a high rate such that the postcooking temperature rise is greater than 5°C, then determine at what temperature or cooking time samples should be removed from the heat to obtain the target endpoint temperature after the postcooking temperature rise (Figure 2). Appropriate thermocouples and recorders should be used for measuring and recording internal temperature.

Changes in color, interpreted as degree of doneness upon completion of cooking, may be influenced by many factors, including animal maturity, type of muscle, fat content, added ingredients, method of cooking, pH, length of cooking, and internal temperature. Color variation in beef patties, such as “premature browning” or “persistent pinking” when cooked to constant final temperatures, poses a unique problem regarding safe final internal temperatures based on USDA (2025). Because of the possibility of color variation in beef patties, degree of doneness based on cooked color cannot be used as a reliable indicator of the selected temperature endpoint; therefore, the actual temperature should be measured. For poultry, the actual endpoint temperature must be measured.

Cookery Methods

Cookery methods can markedly affect the cooking loss and sensory characteristics of meat. A cooking method should be selected with the same care one would use in selecting any other research method and likely varies depending on the research objective. Questions that should be addressed include the following: (1) Is the cooking method similar to common consumer practice? (2) Can the cooking temperature be controlled in a precise and consistent way? (3) Might the cooking method significantly affect the sensory characteristics and perhaps overshadow treatment effects? (4) Will the cooking method cause excessive cooking losses? Most importantly, (5) Will the cooking method provide results relevant to the research question?

The following guidelines are based on results from numerous scientific meat cookery and palatability studies and on the availability and affordability of cooking equipment, as well as realistic industry practices for processed meats. These recommendations (1) reflect methods commonly tested and used by meat scientists, (2) used cookery methods that are consistent and/or minimize cooking effects so that differences are mainly due to treatment effects, (3) do not greatly extend cooking time or cooking losses, and (4) do not mask flavors or impart unusual flavors. These guidelines are not intended to prevent research on new cooking methods as technology advances or to prevent methods of preparation that might enhance palatability and nutrient retention. These methods have been shown to be accurate, consistent, and commonly used and can provide continuity among research institutions and industrial practices.

The terminology and definitions for cooking methods vary among investigators. These guidelines seek to clarify and standardize terminology used in the literature. Cooking steaks in ovens has been described as “oven roasting,” “modified oven roasting,” or “modified broiling,” often inconsistently. In these guidelines, oven broiling and oven roasting are treated as distinct dry-heat cooking methods for intact muscle cuts, and both methods are defined and recommended for use. Regarding moist-heat cookery, many researchers have reported unacceptable palatability and excessive cooking loss when using braizing (Lowe et al., 1952; McDowell et al., 1982; Luchak et al., 1998; Neely et al., 1999; Jeremiah and Gibson, 2003). Therefore, the guidelines do not include braizing, and it is recommended that further research be conducted to provide the changes necessary for this or a similar cooking procedure to be acceptable.

In the past, open-hearth electric grills were commonly used for broiling. Although these units are still manufactured, they exhibit wide variation in cooking surface temperature (Berry and Dikeman, 1994), resulting in inconsistent surface heating that can affect cooking time and cook yield. Cook surface temperature variation may influence sensory qualities of the product and as cooking conditions are not the same within a sample, variation in sensory qualities may be increased within a sample. Open-hearth grills can be acceptable if variation in surface temperature is defined and controlled as much as possible. Gardner et al. (1996) compared impingement oven and clam-shell cookery and demonstrated that they can be effectively used for cookery in meat research. Clam-shell cookery equipment provides a consistent cooking environment, rapid cooking times, and significantly reduced cooking losses compared to open-hearth grills. However, clam-shell grills have been shown to reduce heat induced Maillard reaction products during cooking compared to open-hearth or flat grills. There are also issues with hot and cold spots on some clam shells, including a hot center of grill and cold edges that produce uneven cooking. It is important to understand the positives and negatives of the specific clam shell in each specific laboratory. Conveyor belt grills can provide a consistent cooking process if conditions are properly controlled. Variation in cut thickness and initial internal temperature affect cook rate. Therefore, measures to provide consistent product thickness and internal product temperature are needed. Cook time should be reported along with final internal temperature to describe the cooking environment. Wheeler et al. (1998) compared belt-grill cooking with open-hearth electric grill cooking. Steaks cooked using the belt grill had greater repeatability for Warner-Bratzler shear force (WBSF) and trained sensory panel tenderness traits. Belt-grill instruments, however, are no longer manufactured in smaller laboratory units, but it should be noted that they are a common industry cooking method.

Unless products are intended to be prepared in a microwave or high-velocity forced-air convection oven, these methods are not recommended due to inconsistencies among ovens and cooked product characteristics. Studies show that high-velocity forced-air convection-oven cookery does not meet most of the criteria for selecting research cooking equipment and is not recommended (Lawrence et al., 2001). High-velocity forced air is commonly used in the food industry for pizza toppings, precooked patties, and other similar products. When used for whole-muscle steaks, the rapid heat transfer from top and bottom, rapid cooking, and differences in cook yield from traditional cooking methods such as flat or serrated grills can influence sensory properties and repeatability. If microwave and high-velocity forced-air convection-oven cookery is the intended cooking method, eliminate as much unintended variability as possible. It is recommended to use high-quality microwaves of the same brand and same wattage. Frequent testing of the output power (wattage) should be performed to ensure consistent microwave performance.

Researchers should use a cooking method that best combines consistency and relevance with the goals of minimal cooking losses, relatively rapid cooking times, consistent cooking temperature, minimal masking of normal flavors, and not imparting abnormal flavors. Lack of attention to details in selecting cookery equipment and controlling the cooking process can lead to variability among datasets and interpretation of data especially when comparing data among institutions. As new cooking methods are used for the preparation of meat products, their use in research should be considered if they meet research cookery requirements.

In selecting appropriate cooking equipment, researchers are encouraged to work closely with manufacturers. Information is not readily available regarding the suitability of one particular type, model, or manufacturer over another. In communicating with equipment manufacturers, it is suggested that information be requested from the company regarding the following: sensitivity, accuracy, and precision of temperature control; variation and frequency in temperature swings; and hot- and cold- spot variation. These items can be included in purchase specifications. Before considerable investment is made, it is advisable to test equipment in the manufacturers’ test facilities or consult someone who has used the equipment in the past. Methods other than those described that are demonstrated to be acceptable may be used, but comparable details on data to be collected and methodology need to be defined. Most importantly, regardless of the cooking method and equipment used to conduct sensory and texture evaluation of meat, sufficient attention must be given to the detail required to ensure that reliable data are obtained.

Roasting

Roasting is a method by which heat is transmitted to the meat by convection, either by normal or forced air, in a closed, preheated oven. As indicated above, forced-air convection-oven cookery is not recommended. For traditional roasting, meat is placed on a rack either in or over a shallow pan to catch drippings. The oven door is closed, and the meat is not turned during cooking. For rotisserie-style roasting, meat is placed on a rotating spit. The oven door is closed, and the meat rotates during cooking. If cooking multiple treatments at one time, ensure that the drippings do not contaminate the other treatments. Roasting procedures are reported in Figure 3.

Figure 3.
Figure 3.
Figure 3.

Cooking procedures for roasting meat.

Broiling

Broiling is cooking meat using direct radiant heat, which is also referred to as a “broiler.” The meat is placed above or below the heat source. The heat usually radiates from one direction, so the meat must be turned during cooking. In these guidelines, this method will be referred to as “broiling.” Many investigators prefer broiling because it closely resembles methodology that is commonly used by consumers. Others prefer the “oven-roasting” technique because the cooking conditions can be more easily controlled. Cross et al. (1979) compared roasting and broiling techniques for beef steaks and indicated that both methods were acceptable and resulted in steaks with similar sensory characteristics. Figures 4 and 5 define broiling and belt-grill cookery procedures for meat, respectively.

Figure 4.
Figure 4.

Cooking procedures for broiling meat.

Figure 5.
Figure 5.

Cooking procedures for belt-grill cooking of meat.

Pan broiling

Pan broiling is a method by which small, thin cuts such as patties or bacon are cooked by direct heat of conduction. Meats are placed in frying pans or electric griddles and are cooked without added fat or water. Meats are turned frequently if needed to prevent excess surface browning or crust formation and to allow for more even cooking. The pan or griddle should not be covered during pan broiling. Fat or drippings should be drained between samples to prevent uneven degrees of doneness, cross-contamination across samples, and prevent lipid oxidation. The cooking procedures are presented in Figure 6.

Figure 6.
Figure 6.

Cooking procedures for pan broiling meat.

Impingement oven

Impingement ovens use air heated by gas or electricity that is forced into a chamber or chambers (in large commercial systems, multiple chambers are used to control the cooking method), and the product is carried via a belt, usually a wire mesh belt, through the chamber. Some systems have humidity controls within each chamber. These systems have replaced belt grills for some institutions.

Humidity, room temperature, steak, chop, filet, or patty thickness, and factors influencing heat transfer within the product affect dwell time that is necessary during impingement cooking to reach a defined internal temperature. Different dwell times are expected across days and within a day, and therefore, dwell time and internal endpoint temperature should be monitored routinely within a day and reestablished each cooking day. Impingement ovens are also frequently used to prepare processed meats, specifically pizza toppings (i.e., pepperoni, sausage crumbles, etc.). When evaluating pizza toppings, consider evaluating independently as well as in application, depending on the research objective/s. Dwell time and chamber temperature should be predetermined based on preliminary testing. During preliminary testing, select the dwell time and chamber temperature that produces the appropriate visual and textural degree of doneness of the pizza topping and/or an entire pizza. Methods are presented in Figure 7 for impingement oven cooking.

Figure 7.
Figure 7.

Procedures for cooking meat in an impingement oven.

Clam-shell grilling

Clam-shell cookery simulates grilling by using either a flat griddle or raised grill-mark heated surface with a lid that closes on top of the sample. The top and bottom surfaces are heated simultaneously. This is a popular method of cooking that has largely replaced open-hearth electric grills. Clam-shell cooking procedures are reported in Figure 8. Clam-shell grills may have temperature variations across both surfaces. Preliminary studies should be conducted to determine surface temperature variation. Also, the amount of time needed for a grill—both top and bottom—to return to the predetermined surface temperature after opening and removing a cut, needs to be defined. For grills that automatically turn off, protocols are needed so that this function does not affect cooking or induce variability in meat samples that is not due to treatment effects.

Figure 8.
Figure 8.

Meat cooking procedures for clam-shell grills.

Baking

Baking uses dry heat in an oven from either an electric or gas heat source without air circulation. While oven temperatures can range from 90 to 250°C, meat is usually baked at 175°C. During baking, meat products are either covered or wrapped in foil to provide a moist environment. Baking is commonly used for poultry, fish, and further-processed meat products. Procedures for baking are presented in Figure 9.

Figure 9.
Figure 9.

Meat cooking procedure for baking.

Deep frying

Deep frying involves submerging food in heated shortening or oil. Deep frying is a popular method for cooking breaded meat items. Deep frying, especially for food service products, should be carried out using commercial frying equipment; however, retail frying equipment can be used if necessary. Frying time, oil temperature, oil type, and product load should be predetermined based on preliminary testing. This will vary depending on the product and individual fryer. The oil should be filtered and changed according to oil quality measures to minimize off-flavor migration from the cooking oil. Depending on the treatment, researchers should consider variability between fryers that may contribute to significant sources of variation. During preliminary testing, determine settings that will result in product that meets the required internal temperature and desired breading characteristics. Procedures for deep-frying cooking are presented in Figure 10.

Figure 10.
Figure 10.

Deep-frying cooking procedures for meat.

Pan frying

Pan frying uses a pan placed on a heat source or an electric skillet where a standard amount of fat or oil is added to aid in heat transfer, reduce the adhesion of meat to the pan, and to impart a minimal amount of flavor. Pan-fried meat products are cooked by direct heat of conduction and heat of the added fat. The pan is preheated to a standard temperature, and the preweighed or measured fat or oil is added. Meats are then placed in frying pans or electric skillets. Internal temperature is monitored, as previously defined, if the meat product is of sufficient thickness. If the product is too thin to insert a thermocouple, preliminary studies can be conducted to determine time of cooking. Meat products should be turned over one time during cooking at approximately the midpoint of the cooking process. The pan or skillet should not be covered during pan frying. Fat or drippings should be drained and cleaned from the pan between samples to prevent cross-contamination of flavors across samples and to assure that the fat source is fresh and the same for each product that is cooked. Pan-frying procedures are presented in Figure 11.

Figure 11.
Figure 11.

Cooking procedure for pan frying and sautéing meat.

Sauté

Sautéing includes the addition of flour to the product prior to cooking and uses cooking methods similar to pan frying. As indicated in Figure 11, sautéed meat is dusted or dredged in flour prior to being placed in a pan containing a standard amount of oil. The addition of flour provides a coating during cooking and may affect flavor and cooking parameters of the meat. The weight of the sample before and after the addition of flour is needed to account for flour pickup. The remaining steps for sautéing are the same as those listed for pan frying.

Steeping

Some meat products, such as frankfurters, should be reheated by steeping. A standard amount of water is placed in a pan, the pan is placed on a heating unit, and the water is heated to a standard temperature, which is usually boiling. The product is then placed in the water with the water covering the product. The pan and product are removed from the heat source, covered on top, and allowed to sit in the heated water for a defined amount of time. The amount of water to be added to the pan and the product load should be defined in preliminary studies. The final cook temperature is determined at the end of the steeping time. The steeping time should be reflective of recommended times defined by product manufacturers. Figure 12 defines procedures for steeping meat products. It is important to clarify that steeping and/or boiling should not be used in the evaluation of fresh meat products.

Figure 12.
Figure 12.

Procedures for steeping meat products.

Sous-vide

Sous-vide is a moist cooking method where the product is placed in a heat-stable vacuum-packaged bag and cooked in water (Hrdina-Dubsky, 1989; Church and Parsons, 1993; Creed and Reeve, 1998). Immersion water-circulating systems are used to maintain a predetermined standardized water temperature. Water is circulated by the system to maintain constant water temperature and even temperature exposure by the vacuum-packaged meat sample. The size of the water container and product load per cooking cycle need to be standardized. Water cook temperatures are lower than cook temperatures used in the aforementioned methods and usually are 80°C or less, depending on the product application. Cooking and endpoint temperatures vary in industrial applications across products. Researchers are encouraged to define and emulate industry temperatures and practices. Meat products can be fully or partially cooked using this system. Preliminary studies to determine water temperature, circulation rate, approximate length of cooking time, and internal endpoint temperature need to be determined. Internal product temperature can be monitored by placing a thermocouple in the geometric center of the piece of meat by piercing and threading it through the surface of the vacuum-packaged bag using a preformed, water-proof septum or tape. After cooking, the product is rapidly cooled and then reheated for serving to eat after a period of chilled storage at 1 to 4°C (Creed and Reeve, 1998).

This cooking process reduces cook loss and affects product flavor and texture (Vaudagna et al., 2002; Cabral, 2019; Latoch et al., 2023). With this cooking method, surface dehydration and development of the Maillard reaction during cooking is very limited or does not occur. With the reduced presence of Maillard reaction products and surface browning, sous-vide products are often subjected to a rapid surface heat treatment to obtain brown color on the surface and to allow for the development of Maillard reaction products. The rapid heat treatment can be roasting, broiling, pan broiling, pan frying, or any of the other aforementioned methods that use high heat. Parameters of the reheating process and final endpoint temperature should be recorded. It should be noted that when water temperatures below 62.8°C are used, there are concerns with inactivation of bacteria (Hunt et al., 2023). The texture of sous-vide meat cooked meat differs from that of other cooking methods (Cabral, 2019; Gil et al., 2023; Latoch et al., 2023). Figure 13 defines cooking procedures for sous-vide cooking.

Figure 13.
Figure 13.

Cooking procedures for sous-vide cooking of meat.

Sensory Testing Facilities

Sensory testing environment

The methods used for preparing and presenting samples to the panel require decisions about the testing environment, number of sessions, and physical condition of the samples. The validity of sensory panel results is partially dependent on the control of various factors within the testing environment. Generally recommended procedures for panel room location and layout, lighting, odor control, and comfort of panelists are presented in ASTM MNL26-3RD-EB (2020) and Meilgaard et al. (2025). Detailed recommendations on these items as well as sensory panel booth design can be found in ASTM MNL60 (2016).

The sensory testing facilities should be in an accessible location with sufficient space, temperature and humidity control, and freedom from noise and odors. There should be a slight positive pressure within the sensory room to prevent cooking odors from entering the room. Sufficient electrical outlets and Ethernet or Wi-Fi access should be available for direct data entry. The room/s should be free of other distractions, such as photos or posters on the walls, overly distracting paint colors or wallpaper, or obvious product brand name placement or identification.

Sensory panel booths are commonly used for product evaluations by both trained panels and consumers, as they maximize the ability to control the testing environment (lighting, temperature, food odors, noise, etc.). If space is at a premium, these booths can be portable and capable of quick assembly. Incandescent, fluorescent, or both incandescent and fluorescent lighting can be used in the testing area. Whichever approach is chosen, the lighting should be uniform. The use of a dimmer switch is desirable because it allows for a variety of intensities ranging from 70- to 80-ft candles, which is typical in an office area, to higher levels of 110-ft candles. The light intensity should be measured, reported, and measured across testing days to ensure that uniform lighting conditions are present for each sensory session within a study. If darker lighting is used, differences in appearance may be partially masked or not be as apparent. For any sensory test, visual or oral, light intensity, either foot candles or lux, should be measured and reported.

Color differences may occasionally occur in a study that are not related to the variables being tested and therefore need to be masked. Colored lights (bulbs or filters) can be used in taste panel booths to mask these color differences. Shaded glasses or goggles may also be used; however, Sipos et al. (2021) reported that panelists can easily look around the shaded glasses or move the goggles aside to view the samples. A benefit of adjusted room lighting is a more complete light saturation of the testing environment. Red lights may be most effective in masking variations in reds, pinks, and different degrees of doneness. Amber-colored lights may effectively mask variations in light brown and tan colors. Preliminary studies should be done to determine which light color masks the color differences most effectively. Consider dimming the lights in the sample preparation area so that the light is not transmitted to the booths when the pass-through door is opened.

When testing with consumers, however, colored lights should only be used if absolutely necessary because colored lights can cause artifacts or nontypical responses from consumer panelists. The colored lights might make them suspicious and more attentive to smaller differences in the samples that they would not normally detect using standard lighting. When testing meat products that may be cooked to varying degrees of doneness, consumer panelists whose preferences match the degree of doneness may be preferred. However, if testing the effects of degree of doneness across consumer preferences, understanding consumer degree of doneness preference may assist in interpretation of the results.

Panelists for descriptive analysis panels may be seated around a table, usually a round table, or evaluate individually in booths. Panelists who have extensive training and experience should be able to independently evaluate products using a round table without the biases associated with using untrained panelists. The panel leader should assess the panel’s level of training and only place panelists around a table when they are ensured that this environment will not result in communication and influence of one panelist to others. For consensus data collection, a round table is recommended to allow for discussion.

Additional options for testing locations

While a sensory laboratory is the best location for testing with trained panelists, because of the ability to control the testing environment, there are several alternative options for testing locations when conducting consumer tests. The choice of testing location depends on the test objectives and the compromises that researchers are willing to or must make. The ability to control sample preparation and test administration, which includes noise and distractions, varies greatly depending on the type of testing location that is chosen. Field agency facilities are similar to a sensory laboratory where there is a greater ability to control sample preparation and test administration. To improve testing efficiency, panelists are often seated at individual tables in a classroom-style arrangement with all panelists facing the same direction. School cafeterias and church facilities can also be used, as they provide accessible venues with existing populations for recruitment and allow demographic characteristics to be managed through selection of locations and participant groups within each city. There is less ability to control the atmosphere and distractions in these environments. Testing in school cafeterias, retail stores, mall intercepts, and restaurants have similar issues, including excessive amounts of conflicting aromas, noise, and distractions. The panelist sample size should, therefore, be increased to accommodate the less-controlled testing environment. The number of product samples served should also be minimized as shoppers’ attention is compromised. When using more uncontrolled locations, standard errors would be higher when calculating the recommended sample size when compared to controlled environments. In-home testing is the location that has the least amount of testing control, but it emulates actual use conditions. This is important in testing packaging features, storage, and ease of product preparation.

Consumer testing locations should be vetted and equipped to handle the type of testing required and the products evaluated. Key considerations include adequate kitchen space and the necessary equipment. If the proper equipment is not available, it should be installed. The preparation area should be far enough away from the evaluation area to prevent cross-contamination of odors and the noise distraction. Proper transfer of equipment should be available for bringing prepared samples to the testing area. Storage space (dry, frozen, or refrigerated) for all products, serving equipment, tasting materials, etc. should be available. If using frozen or refrigerated storage, product temperature should be monitored via electronic or manual checks.

Enough staff should be retained to sign-in participants, prepare and serve samples, monitor the testing, and provide payment once testing is complete. The ability to collect and handle data in a robust way is imperative. The facility should have procedures in place for collecting and storing data. If testing in the home, ballots or online surveys should be straightforward and easy to understand, and staff should be available by phone to answer additional questions. The facility should be located near consumers and be able to recruit consumer group/s of interest from the area/s that the product is typically purchased and consumed. For example, consumers of a product not available in the Northwest should not be recruited by a facility in Portland, Oregon.

Preparation and Presentation of Samples to The Panel

Preparation and serving size of sensory samples

Selection of sample preparation method and serving size should be determined based on project objectives and the amount of variation between and within treatments. It is critical that each panelist receives a standardized amount of each sample. If there are multiple preparation methods that consumers typically use in preparing the product being tested, the following factors should be considered:

  • Which preparation method is used most often?

  • Which method yields the best sensitivity and consistent product quality?

  • Which method is more likely to uncover or result in detectable differences?

The cooking method that yields the most sensitivity and consistency may not be the method used most often. If the cooking method results in a less consistent product, the resulting higher degree of variability may cause key sensory differences to be missed due to the loss of test sensitivity. Careful testing should be done to ensure that the cooking method that is selected yields similar sensory properties to the more commonly used cooking method/s. Alternatively, consider including multiple preparation methods if needed.

Regardless of which cooking method is used, it is critical that each panelist receive a standardized amount of each sample. To reduce the chances of introducing variability, samples should be standardized based on weight or dimensions. Serving size should be based on the test objectives, whether the panelists are trained or untrained consumers, and the number of questions or attributes to be evaluated. It is recommended to serve enough of each sample to provide an accurate assessment of the sample and account for inherent variability but not too much to cause taste bud fatigue. Depending on the test objectives, the serving size could be based on the serving size listed on the package label, otherwise use a representative serving size that accounts for the inherent variability in the product. It is recommended to discuss sample preparation and presentation with a sensory professional to develop the best preparation and presentation design of the meat sample.

To account for the moderate to sometimes high degree of variability among and within treatments, meat samples often are cut into cubes, and each panelist receives 2 to 3 cubes from different locations within the piece of meat. For steaks, chops, and roasts, cubes that are 1.27 cm × 1.27 cm × thickness of the cooked cut (2.54 cm) are suggested. If cooking procedures result in variation in cut thickness and charred surfaces, however, the thickness dimension should be standardized and cooked surfaces removed from cuts. Removal of cooked surfaces should be carefully considered, as the influence of cut surface on flavor and texture may be a desirable aspect of the sensory attributes of a sample or may be a component of the sample’s sensory properties. To cut cubes, it is recommended to use a sample sizer. After the hot sample is trimmed of all bone and epimysial (outer) connective tissue, it is placed in the plexiglass sample sizer (Figure 14). The sample sizer should have dimensions of 14-cm long × 12-cm wide × 4-cm deep to accommodate large cuts. On each side, the slots are spaced 1.27 cm apart and have an opening of 3 mm to allow the knife to cut the sample in each direction. For beef patties (depending on the size of beef patties being evaluated [91.5 or 113.5 g]), cooked patties can be cut into 6 or 8 pie-shaped samples as shown in Figure 15. Even with thicker or larger-sized patties, cutting patties into cubes might result in breakage and the inability to obtain equal-sized pieces to serve the panelists. Cutting patties into pie-shaped or wedge samples is recommended.

Figure 14.
Figure 14.

Sensory sample sizer for steaks or chops.

Figure 15.
Figure 15.

Sectioning beef patties for sensory evaluation.

Protocols followed by a trained panel evaluation should be clearly documented to allow for a repeatable and consistent evaluation. During the normal eating experience, consumers assess the juiciness of the meat visually as well as get an initial impression of the tenderness as they cut the sample into bite-sized pieces. Therefore, when conducting consumer tests, it is best to serve samples that are large enough for the panelist to cut to provide a more accurate representation of the actual consumer eating experience. The serving size should be standardized. Location effects within a subprimal should be randomized. The effect of location can then be included in the analysis of variance (ANOVA) model, usually as a random effect.

Depending on the objectives and circumstances of the study, there will be times when it is advisable to serve the smaller 1.27-cm cubes. By randomizing the selection of cubes within a sample, the variability in tenderness within a sample can be better accommodated. These results, however, are a poorer simulation of the normal eating experience when panelists receive bite-sized cubes. Researchers should understand that they are giving up aspects of the eating experience that influence consumer perception when serving 1.27-cm cubes, and therefore, the interpretation and use of the results might be affected. The experimental design might need to be altered (i.e., increase the number of consumers per treatment) if cubes are used.

For poultry, meat and similar products that have been shown to vary significantly by location across the sample, serving size, methods, and evaluation procedures need to be discussed and determined. After preparation, it should be determined what sample size is large enough to account for variability within treatment. For trained panel evaluations, it may be necessary to evaluate poultry pieces independently, such as the head vs. tail portions of the breast portion, thigh vs. drumstick, and so forth. Panelists may be presented with either standardized pieces from each section or a whole chicken breast. If a whole chicken breast is presented, panelists would be instructed to cut the breast in half and eat from the center and remove both the anterior and posterior from the center cut. Panelists would rate individual areas and an overall impression of each attribute. For consumer evaluation, the head and tail ends of the breast should be cut off and discarded. The center portion of the breast should be cut in half and served either randomly, or consumers can receive the same breast portion of the center section of all samples in the study. If samples are randomized by portion, tail or head, these data should be identified and included in the statistical model. Note that with randomizing samples by location, a greater number of consumers may be needed in the study.

Samples for trained panel evaluations for processed meats are typically served by themselves (without condiments or bread). As the objective is to determine if differences are detectable, a greater amount of test sensitivity is desirable. However, when conducting consumer acceptance tests, whether to serve the sample by itself or “in application” (i.e., as a sandwich or on a bun and/or with condiments) should be considered. Serving the sample in application better replicates the actual consumer experience. The risks and benefits of serving in application should be carefully considered, as serving in application introduces variability and may result in fewer differences being detected.

Minimum recommended serving sizes for various processed meat products and factors that should be considered when designing research are presented in Table 2. The study objective, test method, length of questionnaire, and number of samples should guide the actual serving size. It is imperative that sample size be standardized across all servings.

Table 2.

Suggested amount and special factors to consider when serving processed meat products for sensory evaluation

Product Amount Special Factors to Consider
Dinner sausage links Serve 2 ½–3” portions Trim off ends or include?
Frankfurters ½–1 whole link; ¼” coins for descriptive texture analysis Serve with or without bun and condiments?
Breakfast sausage (links) 2 links
Breakfast sausage (patties) 2 patties Discard ends of chub and do not serve exterior slices
Lunch meats (presliced cooked meats) Serve 1–2 whole slices per treatment Discard outer 1–2 slices from package due to color fading and damaged slices
Serve with or without bread and condiments?
Bacon 2 slices
Battered/breaded meats 2-3 nuggets, 1 strip, 1 filet, 1 thigh, 1 breast Do not keep prepared samples warm in a humid environment (i.e., wrapped in foil on a tray)
Whole-muscle marinated meats* 2-oz slices
  • If additional sauces, gravy, toppings, etc. are included as a part of the evaluation.

When evaluating fish whole-muscle products, variation within the filet or in the whole sample has to be accounted for within the sensory sampling. As fish may vary in harvest size, the number of samples or experimental units selected to represent a batch, either wild-caught or farm-raised fish, should be considered. As there is most likely more variation in wild-caught fish, batches usually defined as a catch within a fishing day and selection of fish as experimental units within a batch that are representative of commercial size and weight should be defined. Filet thickness can vary, and there may be variation in texture across the filet. Accounting for variation within a whole-muscle sample, such as discussed for chicken breasts, should be included in the experimental design. Seafood products where the whole piece provides a portion, such as shrimp, clams, oysters, and scallops, should be sampled randomly from a batch. Sufficient product should be presented for trained or consumer evaluation to account for variation within the batch and to provide for at least 2 to 3 portions. In studies where panelists are asked to evaluate multiple attributes, serving portions should be sufficient for evaluating all attributes. For processed seafood products, considerations as defined for processed products across other species are similar.

Sample presentation

Standardized presentation procedures need to be followed to minimize the introduction of extraneous variability. This includes standardizing the critical test elements such as serving temperature and holding conditions, the number of samples per session, and the order of sample presentation.

Serving temperature and holding conditions

Samples should be served at the most appropriate temperature for the attributes being measured. Serving temperature must be achievable and be able to be maintained consistently yet allow for any sensory differences to be detected. Temperature should be monitored to ensure that the samples are a standard temperature when serving and that the temperature is not too high or low. Variation in temperature of samples when the panelists receive them should be documented. As flavor and texture are affected by temperature and holding time, not only should the serving temperature be consistent, but the holding time of the cooked samples should be consistent as well. For each study, procedures for holding, cutting, and serving should be determined ahead of the study.

ASTM Standard E1871-17 (2017) recommends that meat and poultry samples should be served at a minimum temperature of 60°C, while cold products such as deli meats and cold cuts should be stored and served between 2 and 5°C. Under most conditions, the minimum serving temperature of 60°C for warm samples should be adhered to as the maximum volatile aromatics will be detected at this temperature. However, for consumer studies on whole-muscle meats designed with endpoint temperatures to meet the consumer preferences, lower serving temperatures may be needed. Under these conditions, it is recommended that meat be held at 49°C for no more than 20 min in a glass, covered container so that the sensory properties of the steaks and roasts are not affected. If samples are at 49°C when cut, placed in the serving container, and served, the sample should be about 38°C. To humans, the sample will be warm but not cold. Samples should be held as a whole sample and portioned prior to serving to minimize sensory changes due to holding.

When cooking larger cuts of whole-muscle meats to constant temperature endpoints, appropriate time spacing between samples may be difficult to achieve due to different rates of cooking. Some samples may need to be held in a warm environment. Panelists should be informed so that they understand that time between samples (a minimum time should be defined and adhered to) may vary and be longer than the minimum.

Ideally, the samples will be served immediately after being prepared. For fresh meats, this includes serving the samples immediately after being cut, with panelists receiving cubes from various locations within the steak or, in the case of larger portions for consumer panelists, the same steak portion for all treatments. If samples need to be held prior to serving, such as for foodservice types of products, preliminary tests need to be conducted to ensure that holding methods do not affect color, flavor, or texture of the samples. Several procedures for maintaining the temperature of the samples have been used, which include the use of the following sections.

Breaded products

Do not cover or store in an enclosed environment without humidity controls and ventilation, as the breading will become soggy due to moisture migration. It is recommended to keep samples warm using a heat lamp or in a heated deli display case with humidity control (if available).

Fresh meat and nonbreaded products, served warm

Covered pans or glass dishes that are placed in a preheated container of sand, on a warming plate, or heated oven set at a standard temperature, such as 49°C, provide methods of maintaining temperatures prior to and during serving. The temperature of the product during holding, cutting, and serving should be monitored.

Preliminary tests should be conducted to determine the consistency of environmental and product temperatures. The goal is to serve a sample to every panelist that is of the same temperature and that the holding conditions have not influenced the products’ visual, aroma, flavor, basic tastes, and texture sensory attributes. The length of time that samples are held after cooking and prior to serving should be determined and validated to ensure that the sensory characteristics of the product are not altered or that variation is not induced. Preliminary tests should be conducted to validate procedures, and procedures should be routinely checked for compliance. Examples of methods used in meat products include the following: the use of double boilers on an electric hot plate, wrapping the sample in aluminum foil or placing the sample in a glass baking dish with a lid and holding in a heated 49°C oven, or using a preheated yogurt-maker or similar apparatus where samples are stored in glass dishes. Note that storage containers or materials should be odor-neutral and not impart odor or flavor aromatics to the product. Suggested materials are glass containers or aluminum foil.

Lunch meats/deli meats

These products should be held at 2 to 5°C and tempered the appropriate amount of time to achieve a consistent serving temperature. Temperature should be monitored to ensure microbial safety, and products should not be served above 7°C. Samples should be covered to reduce moisture loss or sensory changes during holding prior to sensory evaluation.

Bacon, sausage links, and patties

It is recommended to serve bacon and sausage links and patties immediately after preparation because texture changes very quickly. Sausage and patties can be covered with aluminum foil and held in a steam tray or warming oven usually for no more than 15 min. After 15 min, products may have changes in sensory properties.

Serving containers and utensils

Serving containers and utensils must be made of inert materials that are odor free and nonreactive so that they do not impart any odors or flavors to the samples (ASTM Standard E1871-17, 2017). Serving containers must also be able to maintain sample temperature and moisture content and should be neutral in color unless tint is needed to mask color differences. Glass or white-glazed china (white-glazed porcelain) that has been washed in an unscented detergent, followed by baking at 93°C for several hours, is best but may not be practical for large studies. In these cases, plastic containers and utensils can be used. It should be noted, however, that some plastic materials are less inert, more susceptible to temperature changes, and less odor free than others. All plastic containers should be pretested prior to their use to ensure that under the conditions of use they do not impart sensory changes to the product. The use of wooden toothpicks when serving cubed samples could transfer flavors to the meat and disrupt textural integrity of the sample and should be avoided when possible.

It is recommended that glass containers, such as glass custard dishes, plates, bowls, and other glass containers with lids or concave watch glass covers be used, especially when aroma compounds are to be evaluated. The lids will also keep the samples from drying out while being held. If lids contain material other than glass, preliminary tests should be conducted to make sure the lids do not impact the sensory properties of the samples.

Samples should be coded with 3-digit random numbers. Materials used to mark containers with 3-digit random numbers should be odor free, such as grease pencils, low-odor markers, or labels with printed numbers. If Sharpie® pens are used, the codes should be placed on the containers 24 h prior to use, and the container should sit uncovered to air for 24 h prior to use. Avoid an alphabetical or numerical order to minimize bias.

Number of samples per session and order of presentation

The number of samples that should be presented in a given session is a function of the following: test method—discrimination panel vs. descriptive panel vs. consumer; product characteristics—flavor intensity and spiciness; level of panelist expertise and/or experience; sensory and mental fatigue; number of attributes and modalities (flavor and texture vs. just texture only) to be measured per sample; and use of palate cleansers.

Trained panel evaluations

The number of samples in a discrimination test is dictated by the test method (i.e., triangle vs. duo-trio). Conversely, descriptive analysis tests involve a much larger number of attributes and can include various modalities such as appearance, aroma, flavor, and texture. The number of samples per session in a descriptive analysis test is dictated by the number of attributes and the modality. Taste bud fatigue is a limiting factor for flavor evaluations but less of a concern for texture evaluations. It is recommended to limit the number of samples to 4 to 5 per 1-h session when evaluating the flavor of samples that are not highly seasoned and the ballot length is reasonable. Fewer samples should be evaluated per 1-h session when samples are highly seasoned and/or the ballot is very long. The number of samples that can be accurately evaluated in a session increases with the level of experience and/or training of the panelists. Bohnenkamp and Berry (1987) discussed the effects of sample numbers/session and sessions/day on panelist performance when evaluating beef patties and provided an excellent framework.

Preliminary studies should be conducted to determine the maximum number of samples that can be reasonably evaluated per session without loss of sensory acuity or mental fatigue. Palate cleansers such as room-temperature distilled water and unsalted crackers should also be used to minimize taste bud fatigue and flavor carryover. Additionally, nonfat ricotta cheese is an effective palate cleanser for spicy products or products with metallic, bitter, and lingering aftertastes. Warm water, apple slices followed by plain bagels, or seltzer water can be effective palate cleansers for high-fat products. Adequate time between samples is very important for taste bud recovery. For texture studies, palate cleansers may differ due to the need to remove feeling factors and residual factors in the mouth or nose. Apple slices are good palate cleansers to remove fat mouth-feeling factors, and astringency can be removed using saltless saltine crackers and distilled water. For nasal cavity feeling factors, warm wash cloths that have no residual aromatics can be sniffed. Preliminary studies should be conducted with various palate cleansers to ensure that the palate cleansers are effective and do not affect sensory perception. Time after application of palette cleansers to allow the sensory organs to return to their original state also needs to be defined. After cleansing and the amount of time defined between samples, the palate should be evaluated to ensure no residual or carryover flavors are present (ASTM Standard E1871-17, 2017). Unclean palates contribute to taste bud fatigue and can influence detection and intensity of attributes within the product.

For trained descriptive panels, first sample bias and variation associated with evaluations conducted on different days can be an issue. A standardized warm-up sample, usually a sample that represents the typical product being evaluated in the study, should be served to the panelists at the initiation of the sensory session. After independent evaluation, panelists should discuss and come to consensus for each attribute evaluated. This warm-up sample will assist in standardization among panelists, improve panelists’ concentration, increase panelists’ confidence, and increase calibration of panelists before initiating the sensory session. The sensory leader can use the warm-up sample as a tool to address panel drift and lack of motivation, to increase panelist confidence, and to remove prior environmental factors that can influence the sensory verdict on a given day. This can be done by inserting a warm-up sample from a different treatment so that sensory attribute intensities vary. The panel leader should encourage all panelists to provide input. A panelist who rated an attribute differently should be encouraged to speak out and defend their perceptions. Warm-up samples should be available for additional evaluation during consensus building for confirmation of attribute intensities. The panel leader must ascertain if a panelist is agreeing to an intensity and re-anchoring their perception or agreeing for convenience during consensus building. Panelists should be provided with attribute references to help calibrate and standardize them each sensory day. If variation is apparent during warm-up sessions, panelists should utilize references to anchor intensities.

Consumer panel evaluations

Great care should be taken in setting the number of samples to be served per session in consumer testing. Consumers can differ greatly in their food preferences; therefore, panelists are the greatest source of variability in the statistical model. It is best for all panelists to evaluate all treatments if possible. To minimize taste bud fatigue and loss of interest or concentration among the panelists, the number of samples served per session should be limited based on the product type and the number of questions on the ballot. Palate cleaners as described above can be used to minimize taste bud fatigue and flavor carry over. For unseasoned meat products, consumer panelists can easily evaluate 5 to 6 samples per 1-h session. Spicy or highly seasoned products, such as marinated pork tenderloin with a spicy barbecue flavor, should be limited to 2 to 5 samples per 1-h session.

In situations where the number of samples in the test is greater than what each person can evaluate in one session, it is best to conduct the test over multiple days, with each panelist evaluating a subset of samples each day. If multiday sessions are not an option, then a partially balanced, incomplete-block design can be used, but the overall number of panelists should be increased to achieve the desired number of observations per sample. Consult a statistical text for tables on partially balanced, incomplete-block designs. These designs ensure that each sample or treatment appears with each sample or treatment an equal number of times in any given session.

The length of the questionnaire should also be considered when determining the maximum number of samples to serve and the session length. It is a good practice to conduct preliminary tests with the ballot and the test products to verify that the panelists will be able to complete the test without loss of sensory acuity. Consumer testing generally results in large “halo effects,” so ballots should be as short as possible (Meilgaard et al., 2025). A halo effect is when more than one sensory attribute is evaluated, ratings of each will tend to influence one another.

Every sample should be served in each serving order an equal number of times to reduce any bias related to serving position. Furthermore, every sample should be served before and after every other sample an equal number of times to nullify any bias related to carryover effects. William’s square designs are 1-way to achieve both types of balance (Williams, 1949). For some studies with a large number of treatments, this may not be possible. Order should be randomized using a random number generator or table, and order should be analyzed as a random effect in the model.

Sensory panel participants’ informed consent

The Belmont Report entitled “Ethical principles and guidelines for the protection of human subjects of research” was created by the Department of Health, Education, and Welfare in 1979 to safeguard human subjects used for research, including sensory panels (Department of Health, Education, and Welfare, 1979). This report protects human subjects used in research by fulfilling 3 fundamental ethical principles of respect for persons, beneficence, and justice. These principles protect the autonomy of participants, ensure that they are treated with courtesy and respect, and provide for informed consent. The report maximizes benefits for the research project while minimizing risks to the research subjects and ensures that reasonable, nonexploitative, and well-considered procedures are administered fairly.

In the United States, federally funded research and research conducted at state and federally supported institutions involving human subjects such as sensory panels must obey the ethical rules of the report. Therefore, participants’ informed consent must be obtained, and supervision of the research must be conducted by an institutional review board (IRB) unless the procedures are exempt due to the service nature of the work and not publishing research results. The form prepared for obtaining participants’ consent should precisely convey all the information about the sensory panel to the participants. The consent form should include the purpose of study and study design, who can participate, who will be conducting the research, what participants will be asked to do, possible risks and discomforts, possible benefits/compensation, and who to see with problems or concerns. Figure 16 shows an example of a sensory panel informed-consent form that must be adjusted specifically for each study. For research, the IRB number of the approved procedures should be included in the manuscript.

Figure 16.
Figure 16.

Example of an informed-consent form document (Department of Health, Education, and Welfare, 1979).

Sensory Evaluation Methods

Sensory evaluation uses both subjective and objective procedures. Untrained consumer studies are designed to determine the responses of “typical” consumers. However, through the use of sophisticated panel training and method selection, sensory evaluation also can provide accurate and repeatable objective data. Additional details of the various sensory techniques can be found in Stone and Sidel (2004), Meilgaard et al. (2025), and Lawless and Heymann (2010). The Institute of Food Technologists’ (IFT) guidelines for the preparation and review of papers reporting sensory evaluation data (IFT Sensory Evaluation Division, 1995) should be carefully read by researchers planning to publish in the Journal of Food Science. Approval by the institution’s committee for use of human subjects in research will be needed prior to the initiation of any sensory work, and the IRB number will need to be included in scientific publications.

Scaling is an important aspect of sensory testing methods. Scaling uses words or numbers to express the intensity or degree of an attribute. For many research projects in which sensory traits of meat are evaluated, rating scales are most appropriate. The following types of scales may be used:

  • Graphic or line scales: either a simple line or one marked off into segments can be used. Intensity or degree of the characteristic evaluated must be shown on each end of the scale.

  • Verbal scales: a series of briefly written statements can be used, usually with the name of the sensory attribute with appropriate adverbial or adjectival modifiers that are written out in appropriate order.

  • Numerical scales: a series of numbers can be used, ranging from low to high that are understood to represent successive levels of quality or degrees of a characteristic.

Selection of a scale is an important aspect of a sensory study because scales are the instrument for measuring sensory response. It is important that scales are actionable, easy to use, provide an assurance that the sensory perception of the respondent can be captured, and are repeatable when the same stimulus is evoked. Also, scaling methods should have enough points of determination to capture differences when different levels of intensity are provided. For example, a 5-point scale may not provide enough points of determination for measuring consumer liking or trained sensory intensity but does give reliable results when used as a just-about-right (JAR) scale. Too many points of determination can impact the consistent use of scales. For example, when using 10-cm line scales delineated in 1-point differences, respondents can segment and provide a consistent rating of their sensory stimulus. Data are usually recorded to one unit, such as 1, 3, 5, or 9. However, if the same scale is used as a 100-point scale and the data are recorded as a whole number (91, 48, 14, etc.), respondents may not mark the scale to the nearest one unit, as there are too many points of determination and their scale use may not be sensitive to the whole-number difference. Scales appropriate for each sensory method will be discussed further within each method. Additional guidance on scale selection and use can be found in ASTM Standard E3041-17 (2017).

Discrimination methods

Discrimination or difference testing methods are used to determine when there is a detectable difference between samples. Trained or untrained panelists can be used, depending on the test objectives. If consumers or untrained panelists are used, it is understood that consumers are not as sensitive to differences between products. When trained panelists are used, smaller differences may be discernible.

Discrimination testing methods are divided into 2 broad categories: (1) testing for overall difference and (2) testing for differences in specific attributes. Sensory professionals select the test based on the test objectives. A brief description of the 5 most commonly used discrimination tests used in meat research is included below. Other methods are used and discussed in Meilgaard et al. (2025), Stone and Sidel (2004), and Lawless and Heymann (2010).

The number of panelists required for a given test should be determined prior to conducting the test, regardless the test method used. The recommended number of panelists is based on the test objectives (similarity vs. difference testing), the desired level of test sensitivity, and the amount of variability within samples. Test sensitivity is a function of the levels of α and β, and the levels of either Pd (triangle and duo-trio tests) or δ (duo-trio and tetrad tests). α is the probability of committing a Type I error, which concludes that a perceivable difference exists when, in reality, there is no difference. β is the probability of committing a Type II error, which concludes that no perceivable difference exists when, in reality, there is a perceivable difference. The proportion of discriminators, or Pd, is the proportion of the population represented by the assessors (panelists) that can discriminate between the 2 products in the test. The δ is the effect of size or maximum allowable sensory difference to be detected in the test.

The following guidelines can be used in setting the levels for these factors (ASTM Standard E1885-18, 2018; ASTM Standard E2610-18, 2018; ASTM Standard E3009-24, 2024; Meilgaard et al., 2025). For α, a statistically significant result at 10 to 5% (0.10–0.05) indicates “slight” evidence that a difference is apparent, 5 to 1% (0.05–0.01) indicates “moderate” evidence, 1 to 0.1% (0.01–0.001) indicates “strong” evidence, and less than 0.1% (<0.001) indicates “very strong” evidence. For β, the strength of evidence that a difference is apparent is assessed using the same criteria as above but substituting “is not apparent” for “is apparent.”

When conducting triangle and duo-trio tests, Pd, the maximum allowable proportion of discriminators, falls into 3 categories of Pd being less than 25%, representing small values; Pd being greater than 25% but less than 35%, which represents medium-sized values; and Pd being greater than 35%, representing large values. For duo-trio and tetrad tests, δ, the maximum allowable sensory difference, falls into 3 categories where δ is less than 0.5, representing small values; δ is greater than 0.5 but less than 1, representing medium-sized values; and δ is greater than 1.0, which represents large values.

The actual settings for α and β should be dictated by the level of risk of being wrong in the outcome of the test. In food companies, these risk settings are influenced by the tier level or volume of the brand or product, the degree of product change, and the known sensory variability of the product. α, β, and Pd or δ may be adjusted based on these criteria and historical company practices. Consult a statistician and/or ASTM Standard E1885-18 (2018), ASTM Standard E2610-18 (2018), or ASTM Standard E3009-24 (2024) for recommendations.

Triangle tests

The triangle test is an overall difference test and is believed to be very sensitive. The test does not, however, give directional or attribute information. There is a guessing rate of 33% that is accounted for in the statistical tables, used to determine if a difference exists or not. To conduct the test, panelists are presented with samples either one at a time or all at once. The panelist is asked to evaluate each sample in the specified order and to identify the sample that is different. The instructions define that 2 samples are the same and 1 is different (Figure 17). Each sample should be presented with a 3-digit code, and order should be randomized for each panelist. There are 6 potential serving orders where A and B represent the 2 products: AAB, ABA, BAA, ABB, BAB, and BBA.

Figure 17.
Figure 17.

Sample ballot for triangle testing. Adapted from Meilgaard et al. (2025).

Table 19.7. from Meilgaard et al. (2016) can be used to determine how many panelists are needed based on the predefined levels for α, β, and Pd. If the study objective is to determine whether a detectable difference exists between treatments, the value selected for α is typically smaller than the value selected for β. For example, with α equal to 0.05, β equal to 0.20, and Pd equal to 30%, 40 panelists would be required. If the study objective is to determine if samples are sufficiently similar to be used interchangeably, however, the value selected for β is typically smaller than the value selected for α. For example, with an α equal to 0.20, β equal to 0.05, and Pd equal to 30%, 39 panelists would be required. Results are analyzed by comparing the number of correct responses to the total number presented using the predetermined α, such as in Table 19.8 in Meilgaard et al. (2016). Additional information on how to conduct this test can also be found in ASTM Standard E1885-18 (2018).

Duo-trio tests

The duo-trio test is an overall difference test and is easier for consumers to understand than the triangle test. There is a 50% guessing rate with the duo-trio test, so it is statistically less efficient than the triangle test. Researchers, therefore, need more respondents to conduct this test than the triangle test. Panelists are served a control. After evaluation of the control, panelists are served 2 samples: one is the same as the control and the other is different. They are asked to identify the sample that is the same as the control (Figure 18). Again, directional and attribute information is not obtained. The levels of α, β, and Pd for the test are predetermined based on whether one is testing for difference (α is set smaller than β) or similarity (α is set larger than β). This determines the number of respondents needed as shown in Table 19.9 of Meilgaard et al. (2016). The recommended number of panelists is 69, with α equal to 0.05, β equal to 0.20, and Pd equal to 30% (testing for difference), and 68, with α equal 0.20, β equal 0.05, and Pd equal to 30% (testing for similarity). The control sample rotates between the 2 samples and the order of presentation of the 2 samples rotates as well. Based on this, the serving order with the first sample defined as the control is ABA, AAB, BAB, and BBA. After completion of the test, the total number of responses and the number of correct responses are counted. Using the defined α, the number of correct responses needed to determine a difference is identified in Table 19.10 of Meilgaard et al. (2016). Additional information on how to conduct this test can also be found in ASTM Standard E2610-18 (2018).

Figure 18.
Figure 18.

Sample ballot for duo-trio testing. Adapted from Meilgaard et al. (2025).

Degree of difference

This method also know as Difference-from-Control can be used to determine not only if a detectable difference exists but also the size of differences among a set of samples. Difference-from-control (DFC) testing is also recommended when products are highly variable. This method can be expanded to measure not only the degree of overall difference but also differences in multiple attributes, such as beef-flavor strength or tenderness.

In DFC testing, panelists receive the control sample twice, one time as a labeled reference sample and a second time served blind as a coded sample along with the other coded test samples. Panelists are instructed to taste the reference first, then each coded sample, and indicate the degree of difference from the reference sample using a scale from “none” to “very large difference,” as shown in Figure 19. The scales are typically 7 to 10 points and can be line scales or category scales that are fully anchored or anchored only at the ends (0 = no difference and 7 = very large difference). Twenty to 50 panelists are generally recommended (Meilgaard et al., 2016). The number of panelists required is determined by the amount of variability within samples and the test objective. If the test objective is to detect a difference, the emphasis will be on Type I error. If the objective is to test for similarity, the emphasis will be on Type II error. Testing for similarity will require more sample presentations (panelists) than testing for difference.

Figure 19.
Figure 19.

Sample ballot for difference-from-control testing. Adapted from Meilgaard et al. (2025).

This method can also easily accommodate multiple treatments in a study. With multiple treatments, it is best to present all samples at the same time with clear instructions on the order in which samples are to be evaluated. If, however, it is not possible to present all samples at the same time due to issues such as taste bud fatigue, present only one test sample at a time along with the reference.

Results are analyzed by calculating the mean difference between the identified control and each test sample vs. the difference between the identified control and the blind control. The mean differences are then compared using ANOVA if there are multiple test samples in the study or a paired t-test if there is only one test sample. Additional details on this method and data analysis procedures can be found in Meilgaard et al. (2016).

Tetrad tests

This method can be used to determine whether a perceptible sensory difference exists between 2 samples. The samples may differ in a single sensory attribute or in several attributes. In a tetrad test, samples are served all at the same time if possible. Panelists are presented with 4 coded samples, where 2 samples are from 1 product and the other 2 are from the other product being tested. The panelists are asked to taste the samples in the order indicated and group the 4 samples into 2 pairs based on the level of similarity between samples (Figure 20).

Figure 20.
Figure 20.

Sample ballot for tetrad test. Adapted from ASTM E3009 (2024) and Meilgaard et al. (2025).

There are 6 potential serving orders, where A and B represent the 2 products: AABB, ABAB, ABBA, BBAA, BABA, and BAAB. The recommended number of panelists is based on test objective (testing for difference or similarity) and predefined α, β, and δ. Table 1.1 in ASTM Standard E3009-24 (2024) can be used when testing for difference, and Table 1.2 can be used when testing for similarity. When testing for differences, the recommended number of panelists is 30, with α equal to 0.05, β equal to 0.20, and δ equal to 1.25. The recommended number of panelists when testing for similarity is also 30, with α equal to 0.20, β equal to 0.05, and δ equal to 1.25. The results are tallied, and the level of significance determined by referring to Table A1.3 in ASTM Standard E3009-24 (2024) if testing for a difference and Table A1.4 if testing for similarity.

Because of the nature of the task that the panelists must perform, this method is more efficient statistically than the triangle test or the duo-trio test. However, this method may be less suitable for samples with a high degree of variability or that may cause excessive sensory fatigue or carryover effects. Additional details on this method can be found in ASTM Standard E3009-24 (2024).

Paired comparison tests

Paired comparison tests, also called 2-alternative forced choice tests, are the most common methods used to determine the differences in 1 attribute between 2 samples. The advantage of this method is that it is simple both to administer and for panelists to understand. Two samples are presented to the panelists, and they identify whether the samples are the same or different. There is a 50% guessing rate with this test. Paired comparison tests can either be 1 sided or 2 sided. How the question is asked determines if the test is 1 sided or 2 sided. For a 1-sided test, the question is whether A is more (or less) than B. For 1-sided tests, the α, β, and Pd are determined using the same tables as for duo-trio tests. A 2-sided test asks if A is different from B. When a test is 2-sided, the distribution is bilateral and is also defined as a 2-sided directional difference test. The number of respondents is based on preset levels of α, β, and Pmax and whether the test is 1 sided or 2 sided. Table 19.9 in Meilgaard et al. (2016) can be used to determine the number of respondents needed for 1-sided tests and Table 19.11 for 2-sided tests. The number of correct responses needed for a significant difference is shown in Table 19.10 in Meilgaard et al. (2016) for 1-sided tests and Table 19.12 for 2-sided tests.

Ranking tests

Ranking tests are used to determine differences in several samples for a specific attribute. This test is different from the previously mentioned tests in that this testing method determines how a specific attribute “X” differs among samples. The panelists are provided a set of samples in a balanced, random order and asked to rank the samples from lowest to highest for the intensity of the attribute of interest. It is recommended that all samples be presented at one time, but presentation of samples one at a time also can be done. If respondents are trained, no fewer than 8 respondents should be used. If more than 16 respondents are used, the sensitivity of the test increases. The data are analyzed using the Friedman’s test. Additional information on this method can be found in Meilgaard et al. (2016).

Descriptive analysis methods

Descriptive sensory analysis involves sensory approaches that discriminate and describe both qualitative and quantitative properties using trained panelists. Considerable time and effort are involved in training and maintaining descriptive sensory panels. All sensory attributes of a meat product can be subjected to descriptive analysis or limited to selected criteria such as flavor and texture. The 4 prevalent methods are described below. For a more complete discussion and understanding of descriptive analysis methods, see ASTM MNL13-2ND-EB (2020).

Flavor-profile method

In the flavor-profile method, potential panelists are screened according to their abilities to discriminate aroma and flavor differences. This method uses descriptive terms to characterize the flavor of a product as well as provide the intensities and order of appearance of the various aromas, flavors, and aftertastes detected. After the 4 to 6 members of a panel individually evaluate a sample, the results are submitted to the panel leader who leads a discussion to arrive at a consensus on the sample. Reference samples can be used to present flavor attributes. Data can be presented in tabular, graphic, or verbal format. With this method, there is no accounting for panelist effects.

Texture-profile method

The texture-profile method is based on the same concept as flavor profiling in that the overall texture of a product comprises a number of different texture attributes. Panelists define the procedures and terminology to use in the textural evaluation. Modifications are made in the texture-profile approach as outlined in ASTM MNL13-2ND-EB (2020) compared to the flavor-profile method. For the texture-profile method, more precise definitions and evaluation procedures were developed. New reference scales and various scaling procedures were also defined. With the texture-profile method, there is a collection of individual scores without consensus and statistical analysis of data.

Texture attributes in the texture-profile method can be classified as either mechanical, geometric, or those related to moisture and fat content. Mechanical characteristics are revealed as the meat samples react to stress, such as chewing. Geometric characteristics relate to size, shape, and orientation of the product before and during breakdown. The contributions of moisture and fat are determined through mouthfeel. Texture references have been developed and can be used for training. An example of a texture lexicon for ground beef patties was published by AMSA (1983). A lexicon is defined as a group of attributes, their definition, and references within a product or product type. An example ballot that includes flavor and texture attributes for ground beef patties is presented in Figure 21. This ballot uses beef-flavor attributes from the beef-flavor lexicon (Adhikari et al., 2011), as discussed below, and texture attributes from AMSA (1983). The texture attributes can be easily defined and scaled using solid oral texture attributes from Meilgaard et al. (2016). Note that not all texture attributes for ground beef were included on the ballot—only those important to the hypothesis of the study were included.

Figure 21.
Figure 21.

Example of a flavor and texture descriptive attribute ballot for ground beef patties based on AMSA (1983), Meilgaard et al. (2016), and Miller et al. (2019); panelists provide ratings for each attribute from each sample provided.

Quantitative-descriptive analysis

Quantitative-descriptive analysis (QDA) was developed to provide a stronger statistical treatment of data than data from profile-type methods. While training is provided regarding methods and terminology, to some degree panelists are free to develop their own approach to scoring. Overall discussions of the evaluations after a session are not usually conducted. Data are reported in the form of a spiderweb with a branch or spoke for each attribute. An example of this form of data reporting is shown in Figure 22. Additional information on this method can be found in ASTM MNL13-2ND-EB (2020).

Figure 22.
Figure 22.

Spider web from quantitative-descriptive analysis wherein 3 levels of sodium lactate were added to pork roasts.

SpectrumTM-descriptive attribute analysis method

The Spectrum method (Meilgaard et al., 2025) is a custom-designed approach, providing detailed information on the sensory attributes of a product including aroma, flavor, and texture, as well as their intensities using absolute or universal scales. A lexicon, or dictionary of attributes, references, and examples, is used as a basis to present specific attributes of a product. With this method, line scales or the universal 16-point intensity scale is used. Munoz and Civille (1998) discuss the use of the universal scale and lexicon development. Also, extensive panelist training is conducted to ensure that each panelist understands each attribute, can scale each attribute for intensity, and can consistently accomplish this task across products and over time. Attribute definitions and references provide a clear understanding of the perception and scaling of an attribute. It also provides an avenue to reduce variation across panelists and sensory days.

With this method, data are analyzed using ANOVA to determine differences across products within an attribute. Additionally, multivariate analyses can be used to understand how multiple attributes relate to a treatment or product. Any sensory category of attributes (visual appearance, odor, flavor aromatics, basic tastes, mouthfeel, texture, and aftertastes) are multivariate, and individual attributes that are components of the multivariate response are univariate (i.e., sour, salty, beef identity, etc.). Analyses of these data are presented in the statistical analysis section.

Short-version Spectrum descriptive method for quality assurance and shelf-life studies

This method is based on the Spectrum method and applies the same principles when a full lexicon of attributes is not necessary. The sensory professional determines, either through lexicon development or product knowledge, the main attributes that need to be evaluated for a product. These attributes and references are defined, panelists are trained using references and scaling exercises, and the products are evaluated. See Munoz et al. (1992) for a more comprehensive understanding of how to apply this method.

Meat-descriptive attribute evaluation

In the first guidelines in 1978, a unified method of evaluating meat palatability attributes was introduced (AMSA, 1978). This method concentrated on juiciness, connective tissue amount, muscle-fiber tenderness, and overall tenderness as the major attributes of meat palatability. The 2 previous versions of the guidelines also included overall flavor intensity. This attribute is a general intensity attribute and does not describe a specific flavor attribute. For example, 2 samples may have the same intensity of overall flavor; however, one sample may have predominantly beef identity and browned flavor aromatics, and the other sample may have predominantly liver-like and putrid flavor aromatics. Also, it is difficult to reference and scale across panels. When possible, it should be replaced with the species-specific flavor lexicons, where each flavor attribute can be referenced and scaled. If it is not practical to train for the entire flavor lexicon or if it is not necessary based on the experimental objectives, 1 or 2 flavor notes, such as beef-flavor identity and fat-like flavor aromatics, could be used (Munoz et al., 1992; Meilgaard et al., 2025). Selection of attributes to include in the lexicon should be based on the study objectives. This method uses an 8- or 9-point, verbal anchored scale (Figure 23), or it can be modified to a 15- or 16-point scale to more easily incorporate flavor attributes (Figure 24). Regardless of scale, the method requires training on each attribute, panel performance validation, and use of correct cooking and product sampling techniques.

Figure 23.
Figure 23.

Example ballot for beef descriptive flavor and texture attributes using 15- and 16-point scales (Miller, 2013).

Figure 24.
Figure 24.
Figure 24.
Figure 24.
Figure 24.

Examples of training sessions for a beef-flavor descriptive attribute panel using the beef-flavor lexicon (Adhikari et al., 2011; Miller et al., 2023).

Magnitude estimation

Magnitude estimation involves the assignment of numbers to indicate intensity in relation to the first sample or reference sample. All subsequent samples are evaluated in proportion to the first sample rating. For example, if the sample is perceived as having a salty rating of 1, the sample would be rated 4 if it is 4 times saltier than the reference sample.

Lexicons

Descriptive methods rely on training and definition of attributes. Note that with QDA and magnitude estimation, lexicons may or may not be used. Lexicons are dynamic documents where attributes can be added or removed for individual studies. The lexicons presented use the universal scale, but the scaling of lexicon attributes can be adjusted to be product specific. For example, when examining differences in chicken wing marinades, where the objective is to reduce the level of salt, the universal scale defines 0.2% sodium chloride in water as 2.5 out of 15 and 0.35% sodium chloride in water as 5.0 out of 15. The wing treatments may fall within the defined references, but differences may not be detected because the saltiness is only rated between 2.5 and 5 and the entirety of the scale is not used. For example, the scale can be changed to a product-specific scale, where 0.2% sodium chloride in water is 5 out of 15 and 0.35% sodium chloride in water is 10 out of 15. Extensive panel training is needed to familiarize panelists with the use of the scale. Attribute definitions and references should be documented so that the experiment can be repeated or an understanding of what was measured is clear during the discussion of results.

Prior to conducting a study, if a lexicon is applicable, it should be used as a basis for panel training as discussed below (Tables 313). While these are not all-inclusive, they provide a base for lexicon use. The beef-flavor lexicon developed by Adhikari et al. (2011) and adapted by Miller et al. (2023) is presented in Table 3. This lexicon defines major aroma and flavor attributes found in whole-muscle beef. Table 4 presents the pork lexicon developed by Chu (2015) that can be used similarly to the beef lexicon but for whole-muscle pork. A lamb lexicon, developed and used by Wall et al. (2024), is presented in Table 5 and includes 2 texture attributes in addition to odor, flavor aromatics, basic tastes, and aftertastes. A lexicon for poultry is presented in Table 6, and definitions and references for potential fish descriptive attributes are presented in Table 7 (Johnsen et al., 1987; Farmer et al., 2000; Lazo et al., 2017). For aquatic species, sensory attributes differ greatly. Table 7 attributes are a culmination of multiple studies and are an aggregate of terms. During training, samples from the study and additional samples should be provided to panelists to familiarize and extend their experience. If attributes are not present or attributes are found that are not reported in the lexicon, attributes can be removed or added.

Table 3.

The beef-flavor lexicon defined by Adhikari et al. (2011) and as adapted by Miller et al. (2023)

Attribute Definition References
Animal hair Aromatics perceived when raw wool is saturated with water Caproic acid (hexanoic acid) = 12.0 (A)
Beef identity* Amount of beef-flavor identity in the sample Swanson® beef broth = 5.0 (F)
80% lean ground beef = 7, flavor 0 (A and F)
Beef brisket = 11.0 (A and F)
Bitter* Fundamental taste factor associated with a caffeine solution 0.01% caffeine solution = 2.0 (F)
0.02% caffeine solution = 3.5 (F)
Bloody/serum-like* Aromatics associated with blood on cooked meat products; closely related to metallic aromatic USDA Choice strip steak = 5.5 (A and F)
Beef brisket = 6.0 (aroma and flavor)
Brown/roasted* Round, full aromatic generally associated with beef suet that has been broiled Beef suet = 8.0 (A and F)
80% lean ground beef = 10.0 (A and F)
Burnt Sharp/acrid flavor note associated with over-roasted beef muscle, something over-baked or excessively browned in oil Alf’s Puffed Red Wheat® = 5.0 (A and F)
Chemical Aromatics associated with garden hose, hot Teflon pan, plastic packaging, and petroleum-based products like charcoal lighter fluid Ziploc® sandwich bag = 13.0 (A)
Clorox® in water = 6.5 (F)
Cocoa Aromatics associated with cocoa beans, powdered cocoa, and chocolate bars; brown, sweet, dusty, often bitter aromatics Hershey’s® cocoa powder in water = 3.0 (F)
Hershey’s® chocolate kiss = 7.5 (A), 8.5 (F)
Cooked milk Combination of sweet, brown flavor notes and aromatics associated with heated milk Mini Babybel® original Swiss cheese = 2.5 (F)
Dillon’s whole milk = 4.5 (F)
Dairy Aromatics associated with products made from cow’s milk, such as cream, milk, sour cream, or butter milk Dillon’s reduced fat milk (2%) = 8.0 (F)
Fat-like* Aromatics associated with cooked animal fat Hillshire Farms Beef Lit’l Smokies® = 7.0 (A and F)
Beef suet = 12.0 (A and F)
Green Sharp, slightly pungent aromatics associated with green/plant/vegetable matters such as parsley, spinach, pea pod, fresh-cut grass, etc. Hexanal in propylene glycol (5,000 ppm) = 6.5 (A)
Fresh parsley water = 9.0 (F)
Green-hay Brown/green dusty aromatics associated with dry grasses, hay, dry parsley, and tea leaves Dry parsley in medium snifter = 5.0 (A)
Dry parsley in ∼30-mL cup = 6.0 (F)
Leather Musty, old leather (like old book bindings) 2,3,4-Trimethoxybenzaldehyde = 3.0 (A)
Liver-like Aromatics associated with cooked organ meat/liver Beef liver = 7.5 (A and F)
Braunschweiger liver sausage = 10.0
(A and F—must taste and swallow)
Metallic* Impression of slightly oxidized metal, such as iron, copper, and silver spoons 0.10% potassium chloride solution = 1.5 (F)
USDA Choice strip steak = 4.0 (A and F)
Dole® canned pineapple juice = 6.0 (A and F)
Overall sweet* Combination of sweet taste and sweet aromatics; the aromatics associated with the impression of sweet Post Shredded Wheat®, spoon size = 1.5 (F)
Hillshire Farms Beef Lit’l Smokies® = 3.0 (F)
SAFC ethyl maltol 99% = 4.5 (A)
Rancid Aromatics commonly associated with oxidized fat and oils; may include cardboard, paint-like, varnish, and fishy Microwaved Wesson® vegetable oil (3 min at high) = 7.0 (F)
Microwaved Wesson® vegetable oil (5 min at high) = 9.0 (F)
Salty* Fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5 (F)
0.25% sodium chloride solution = 3.5 (F)
Sour* Fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5 (F)
0.050% citric acid solution = 3.5 (F)
Sour aromatics* Aromatics associated with sour substances Dillon’s buttermilk = 5.0 (F)
Sour dairy Sour, fermented aromatics associated with dairy products such as buttermilk and sour cream Laughing Cow® light Swiss cheese = 3.0 (A), 7.0 (F)
Dillon’s buttermilk = 4.0 (A), 9.0 (F)
Spoiled Presence of inappropriate aromatics and flavors that are commonly associated with the products; a foul taste and/or smell that indicates the product is starting to decay and putrefy Dimethyl disulfide in propylene glycol (10,000 ppm) = 12.0 (A)
Sweet* Fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0 (F)
Umami* Flat, salty, somewhat brothy; taste of glutamate, salts of amino acids, and other molecules called nucleotides 0.035% accent flavor enhancer solution = 7.5 (F)
Warmed-over Perception of a previously cooked and reheated product 80% lean ground beef (reheated) = 6.0 (F)
Smoky charcoal An aromatic associated with meat juices and fat dripping on hot coats which can be acrid, sour, burnt, etc. Wright’s Natural Hickory seasoning in water = 9.0 (A)
Smoky wood Dry, dusty aromatic reminiscent of burning wood Wright’s Natural Hickory seasoning in water = 7.5 (A)
Buttery Sweet, dairy-like aromatic associated with natural butter Land O’Lakes unsalted butter tasted = 7.0 (F and A)
Refrigerator stale Aromatics associated with products left in the refrigerator for an extended period time and absorbing a combination of odors (lack of freshness/flat) 80/20 ground beef, cooked, left chilled overnight = 6.0 (F), 8.0 (A)
Soapy An aromatic commonly found in unscented hand soap 0.12 oz Clorox wipe liquid in 4 oz water= 3.0 (A)
0.5 g Ivory bar soap in 100 mL water = 6.5 (A)
Barnyard Combination of pungent, slightly sour, hay-like aromatics associated with farm animals and the inside of a horn White pepper in water = 4.5 (F)
Tinture of civet = 6.0 (F)
Heated oil The aromatics associated with oil heated to a high temperature Wesson oil, microwaved 3 min = 7.0 (F and A)
Lay’s potato chips = 4.0 (A)
Asparagus The slightly brown, slightly earthy green aromatics associated with cooked green asparagus Asparagus water = 7.5 (A)
Asparagus water = 6.5 (F)
Cumin The aromatics commonly associated with cumin and characterized as dry, pungent, woody and slightly floral McCormick or Shilling ground cumin = 10.0 (A)
McCormick or Shilling ground cumin = 7.0 (F)
Floral Sweet, light, slightly perfume impression associated with flowers 0.12 oz Clorox wipe liquid in 4 oz water = 8.0 (A)
Geraniol = 7.5 (A)
1:1 white grape juice to water = 5.0 (F and A)
Beet A dark damp, musty, earthy note associated with canned red beets Food club sliced beets = 6.0 (F)
Food club sliced beets = 4.0 (F)
1 part juice to 2 parts water served ½ oz in 1 oz cups
Petroleum-like A specific chemical aromatic associated with crude oil and its refined products that have heavy oil characteristics Vaseline petroleum jelly = 3.0 (A)
½ tsp of Vaseline in 1-oz cups
  • Abbreviations: A, aroma; F, flavor.

  • Represents major attributes that should be present at some level in all samples.

Table 4.

The pork-flavor lexicon adapted from Chu (2015)

Attribute Definition Reference
Astringent The chemical feeling factor on the tongue or other skin surfaces of the oral cavity, described as puckering/dry and associated with tannins or alum Lipton tea, 1 bag = 6.0 (F)
Lipton tea, 3 bags = 12.0 (F)
Boar taint Aromatic associated with boar taint; hormone-like; sweat, animal urine 0.1 g 3-methylindole = 13.0 (A)
Androstenone = 15.0 (A)
Bitter The fundamental taste factor associated with a caffeine solution 0.05% caffeine in water = 2.0 (F)
0.08% caffeine in water = 5.0 (F)
Bloody/serum-like An aromatic associated with blood on cooked meat products; closely related to metallic aromatic Boneless pork chop, 135°F = 2.0 (F and A)
Brown/roasted A round, full aromatic generally associated with pork suet that has been broiled Pork fat, cooked = 3.0 (F), 4.0 (A)
Burnt The sharp/acrid flavor note associated with over-roasted pork muscle, something over-baked or excessively browned in oil Arrowhead Puffed Barley Cereal® = 5.0 (A and F)
Cardboardy Aromatic associated with slightly oxidized fats and oils, reminiscent of wet-cardboard packaging Dry cardboard = 5.0 (F), 3.0 (A)
Wet cardboard = 7.0 (F), 6.0 (A)
Chemical Aromatic associated with garden hose, hot Teflon pan, plastic packaging, and petroleum-based products such as charcoal lighter fluid 1 drop Clorox in 200 mL water = 6.5 (F)
Ziploc bag = 13.0 (A)
Fat-like Aromatics associated with cooked animal fat Pork fat, cooked = 10.0 (F); 7.0 (A)
Floral Sweet, light, slightly perfume impression associated with flowers 0.12 oz Clorox wipe liquid in 4 oz water = 8.0 (A)
Geraniol = 7.5 (A)
1:1 white grape juice to water = 5.0 (F and A)
Heated oil The aromatics associated with oil heated to a high temperature Wesson oil, microwaved 3 min = 7.0 (F and A)
Lay’s potato chips = 4.0 (A)
Liver-like Aromatics associated with cooked organ meat/liver Pork liver, cooked = 15.0 (F), 12.0 (A)
Metallic The impression of slightly oxidized metal, such as iron, copper, and silver spoons Dole pineapple juice = 6.0 (A and F)
0.10% KCl solution= 1.5 (A and F)
Nutty Nutty characteristics are as follows: sweet, oily, light brown, slightly musty and/or buttery, earthy, woody, astringent, bitter, etc. Diamond shelled walnut, ground for 1 min = 6.5 (F)
Pork identity Amount of pork-flavor identity in the sample Boneless pork chop, 175°F = 7.0 (F), 5.0 (A)
80/20 ground pork, cooked = 6.0 (F), 5.0 (A)
Refrigerator stale Aromatics associated with products left in the refrigerator for an extended period time and absorbing a combination of odors (lack of freshness/flat) 80/20 ground pork, cooked, left chilled overnight = 6.0 (F), 8.0 (A)
Salty The fundamental taste factor of which sodium chloride is typical 0.2% salt in water = 2.5 (F)
0.35% salt in water = 5.0 (F)
Soapy An aromatic commonly found in unscented hand soap 0.12 oz Clorox wipe liquid in 4 oz water = 3.0 (A)
0.5g Ivory bar soap in 100 mL water = 6.5 (A)
Sour The fundamental taste factor associated with citric acid solution 0.05% citric acid in water = 2.0 (F)
0.08% citric acid in water = 5.0 (F)
Spoiled/putrid The presence of inappropriate aromatics and flavors that are commonly associated products; it is a foul taste and/or smell that indicates product is starting to decay and putrefy Boneless pork chop, 175°F, left out for 24 h, then refrigerate for 6 d = 3.0 (A)
80/20 ground pork, cooked, same as above = 5.0 (A)
Sweet The fundamental taste factor associated with a sucrose solution 0.05% sugar in water = 2.0 (F)
0.08% sugar in water = 5.0 (F)
Umami Flat, salty, somewhat brothy; the taste of glutamate, salts of amino acids, and other molecules called nucleotides 0.035% accent flavor = 7.5 (F)
Vinegary Aroma notes associated with vinegar 1.1 g vinegar in 200 g water = 6.0 (F), 4.0 (A)
Warmed-over Perception of a product that has been previously cooked and reheated 80/20 ground pork, cooked, left chilled overnight, and reheated = 5.0 (A and F)
  • Abbreviations: A, aroma; F, flavor.

Table 5.

Definitions and references for lamb descriptive attributes and scaling intensities adapted from Wall et al. (2024)

Attribute Definition References
Lamb identity Amount of lamb identity in the sample Knorr lamb stock cubes = 3.0
Lean ground lamb = 6.0
Lamb broth sample (from crockpot) = 7.0
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Sweet The fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0
Bitter The fundamental taste factor associated with a caffeine solution 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution= 1.5
0.045% citric acid solution= 3.0
0.065% citric acid solution = 5.0
Umami Flat, salty, somewhat brothy; the taste of glutamate, salts of amino acids, and other molecules called nucleotides 0.035% accent flavor enhancer solution = 7.5
Fat-like Aromatics associated with cooked animal fat Hillshire Farms Lit’l Beef Smokies = 7.0
80% lean ground lamb = 6.0
Lamb suet (seared) = 10.0
Brown A round, full aromatic generally associated with lamb suet that has been broiled Lamb suet (seared) = 7.0
80% lean ground lamb = 4.0
Roasted A round, full aromatic generally associated with lamb that has been broiled/roasted Lamb crock pot roast = 11.0
Metallic The impression of slightly oxidized metal, such as iron, copper, and silver spoons 0.10% potassium chloride solution = 1.5
Lamb leg steak grilled to 58°C = 5.0
Dole canned pineapple juice = 6.0
Liver-like Aromatics associated with cooked organ meat/liver Beef liver (broiled) = 12.0
Braunschweiger liver sausage = 10.0
Cardboardy Aromatic associated with slightly oxidized fats and oils, reminiscent of wet-cardboard packaging Dry cardboard = 5.0
Wet cardboard = 7.0
Bloody/serum-like The aromatics associated with blood on cooked meat products, closely related to metallic aromatic Lamb stew pieces grilled to 63°C = 6.0
Leg steak grilled to 63°C = 3.0
Musty/earthy Musty, sweet, decaying vegetation Sliced button mushrooms = 3.0
Le Nez Du Café No. 1 = 8.0
Lanolin/animal hair The aromatics perceived when raw wool is saturated with water; a sweet, oily aromatic off-flavor associated with lamb Raw wool = 4.0
Bag balm = 9.0
Mutton Strong, musty, gamey aromatics characteristic of meat from older sheep Chopped fresh sage = 2.0
Wool from 2-y-old Angora billy = 6.0
4-methyloctanoic acid (200 μL on cotton ball in snifter) = 10.0
Green Sharp, slightly pungent aromatics associated with green/plant/vegetable matters such as parsley, spinach, pea pod, fresh-cut grass, etc. Fresh parsley water = 9.0
Juiciness The amount of perceived juice that is released from the product during mastication Carrot = 8.5
Mushroom = 10.0
Cucumber = 13.5
Apple = 13.5
Watermelon = 15.0
Beef strip loin grilled to 58°C = 9.0
Beef strip loin grilled to 80°C = 4.0
Muscle-fiber tenderness The ease in which the muscle-fiber fragments during mastication Beef eye of round steak cooked to 70°C = 4.0
Beef strip loin cooked to 70°C = 8.0
Beef tenderloin cooked to 70°C = 14.0
Connective tissue The structural component of the muscle surrounding the tissue amount during mastication Beef brisket steak cooked to 70°C = 4.0
Beef tenderloin cooked to 70°C = 11.0
  • Abbreviation: F, flavor.

Table 6.

Definitions and references for poultry descriptive attributes and scaling intensities adapted from Laird et al. (2024)

Attribute Definition References
Chicken identity Amount of chicken identity in the sample Chicken breast grilled to 71°C = 4.0 (F)
Ground chicken cooked in skillet set at 350°F to 71°C internal temperature = 5.0 (F)
Swanson’s chicken broth = 7.0 (F)
Dark chicken baked thigh to 175°C internal temperature = 6.0 (F)
White chicken breast baked to 175°C internal temperature = 4.0 (F)
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Sweet The fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0
Bitter The fundamental taste factor associated with a caffeine solution 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5
0.045% citric acid solution = 3.0
0.065% citric acid solution = 5.0
Umami Flat, salty, somewhat brothy; the taste of glutamate, salts of amino acids and other molecules called nucleotides 0.035% accent flavor enhancer solution = 7.5
Fat-like Aromatics associated with cooked chicken fat Chicken fat from the thigh, covered with water, cooked in pan with lid, boiled for 20 min, remove lid and cooked until the water evaporates = 8.0 (F)
Grilled chicken skin in skillet set at 350°F until brown = 5.0 (F)
Brown A round, full aromatic generally associated with chicken suet that has been broiled Chicken suet (seared) = 7.0
Roasted A round, full aromatic generally associated with chicken that has been roasted Chicken breast, skin off, from a whole roasted chicken to 165°F
Metallic The impression of slightly oxidized metal, such as iron, copper, and silver spoons 0.10% potassium chloride solution = 1.5
Dole canned pineapple juice = 6.0
Fishy Aromatics associated with fish Canned StarKist tuna = 12 (F), 10 (A)
Canned chicken = 4 (F)
Cardboardy Aromatic associated with slightly oxidized fats and oils, reminiscent of wet-cardboard packaging Dry cardboard = 5.0
Wet cardboard = 7.0
Bloody/serum-like The aromatics associated with blood on cooked meat products, closely related to metallic aromatic Boneless pork chop, 135°F = 2.0 (F and A)
Musty-earthy Musty, sweet, decaying vegetation Sliced button mushrooms = 3.0
Le Nez Du Café No. 1 = 8.0
Chemical Aromatic associated with garden hose, hot Teflon pan, plastic packaging, and petroleum-based products such as charcoal lighter fluid 1 drop Clorox in 200 mL water = 6.5 (F)
Ziploc bag in snifter = 2.0 (A)
Floral Sweet, light, slightly perfume impression associated with flowers (A) 0.12 oz Clorox wipe liquid in 4 oz water = 8.0 (A)
Geraniol, 2 drops on cotton ball in snifter = 7.5 (A)
1:1 white grape juice to water = 5.0 (F and A)
Green Sharp, slightly pungent aromatics associated with green/plant/vegetable matters such as parsley, spinach, pea pod, fresh-cut grass, etc. Fresh parsley water = 9.0
Liver-like Aromatics associated with cooked organ meat/liver Chicken liver 71°C = 9.0 (F)
Heated oil The aromatics associated with oil heated to a high temperature Wesson oil, microwaved 3 min = 7.0 (F and A)
Lay’s potato chips = 4.0 (A)
Nutty Nutty characteristics are as follows: sweet, oily, light brown, slightly musty and/or buttery, earthy, woody, astringent, bitter, etc. Diamond shelled walnut, ground for 1 min= 6.5 (F)
Paint-like Wesson oil placed in covered glass container in 100°C oven for 14 d = 8 (F), 10 (A)
Burnt The sharp/acrid flavor note associated with over-roasted chicken muscle, something over-baked, or excessively browned in oil Arrowhead Barley Cereal, 7–10 puffs = 3.0 (F and A)
Refrigerator stale Aromatics associated with products left in the refrigerator for an extended period time and absorbing a combination of odors (lack of freshness/flat) Ground chicken, 71°C, left chilled overnight, served room temperature = 6.0 (F), 8.0 (A)
Soapy An aromatic commonly found in unscented hand soap 0.12 oz Clorox wipe liquid in 4 oz water = 3.0 (A)
0.5 g Ivory bar soap in 100 mL water = 6.5 (A)
Spoiled/putrid The presence of inappropriate aromatics and flavors that are commonly associated with spoiled products; it is a foul taste and/or smell that indicates spoilage product is starting to decay and putrefy; usually associated with sulphur-based compounds Boneless chicken breast stored at room temperature, raw for 24 h, refrigerated for 6 d, 175°F, smelled only = 8–12 vinegary (A)
Vinegary Aroma notes associated with vinegar 1.1g Vinegar in 200g water = 6.0 (F); 4.0 (A)
Warmed-over Warmed-over perception of a product that has been previously cooked and reheated Boneless, skinless chicken breast, cooked to 71°C, left chilled overnight, and microwaved for 1 min = 5.0 (F and A)
Texture
Juiciness The amount of perceived juice that is released from the product during mastication Carrot = 8.5
Mushroom = 10.0
Cucumber = 13.5
Apple = 13.5
Watermelon = 15.0
Beef strip loin grilled to 58°C = 9.0
Beef strip loin grilled to 80°C = 4.0
Muscle-fiber tenderness The ease in which the muscle-fiber fragments during mastication Beef eye of round steak cooked to 70°C = 4.0
Beef strip loin cooked to 70°C = 8.0
Beef tenderloin cooked to 70°C = 14.0
Connective tissue The structural component of the muscle surrounding the tissue amount during mastication Beef brisket steak cooked to 70°C = 4.0
Beef tenderloin cooked to 70°C = 11.0
  • Abbreviations: A, aroma; F, flavor.

Table 7.

Definitions and references for fish descriptive attributes and scaling intensities adapted from Johnsen et al. (1987), Farmer et al. (2000), and Lazo et al. (2017)

Attribute Definition References
Appearance
Color intensity Color intensity from white to light brown inside the flesh of the fish
Pink (salmon) Intensity of pink color in the uncut steak
Peach (salmon) Intensity of peach color in the uncut steak
Orange (salmon) Intensity of orange color in the uncut steak
Beige Beige
Whiteness Whiteness
Paleness Paleness
Coagulation Amount of coagulation that appears on the surface of the salmon steak
Juicy appearance Amount of juice that has seeped onto the plate from the salmon
Separation Oiliness of the juice seeping out of the salmon
Moist appearance Visible moistness of steak when cut with knife
Flaky Visible flakiness of steak when cut with a knife
Crumbly Crumbliness of steak when cut with a knife
Color uniformity Color homogeneity inside the flesh of the fish without black veins or spots
Exudate quantity Quantity of liquid released after cooking the sample
Fat droplets Fat released in fish exudate in the form of oil droplets
Turbidity of exudate Suspended particles in exudate that block transparency 0.035% accent flavor enhancer solution = 7.5
Odor Aromatics associated with cooked animal fat Hillshire Farms Lit’l Beef Smokies = 7.0
Butter intensity Odor of butanedione
Earthy intensity Odor of humid dirt or potting soil
Sardine intensity The odor of fish oil
Odor or flavor aromatics and basic tastes
Salmon-like Intensity of distinctive salmon-like (A and F)
Oily Intensity of oily and fish oils (A and F)
Stagnant Intensity of any stagnant water (A)
Farmyard Intensity of manure/cow dung (A and F)
Fishy Intensity of any other fish-like flavors Canned StarKist tuna = 12 (F), 10 (A)
Fishy, skin Intensity of nonsalmon fishy flavor near the skin of the salmon steak
Nutty The aromatic associated with kesh pecans and other hardshell nuts
Chicken-like The aromatic associated with sweet, cooked chicken meat
Fat complex The aromatic associated with dairy lipid products, melted vegetables, shortening, and cooked chicken skin
Corn The aromatic associated with cooked corn kernels
Green vegetable/grassy The aromatic associated with fresh grassy vegetation and green vegetables
Eggy/sulphur The aromatic associated with boiled old-egg proteins
Geosmin/dry musty The aromatic associated with an old book Geosmin
Wet musty The aromatic associated with mud 2-methylisoborneol
Decaying vegetation The aromatic associated with decaying vegetation particularly pondweed, decaying wood, and swamp grass
Cardboardy
TMA The aromatic associated with the reference trimethylamine (TMA)
Sweet The fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Bitter The fundamental taste factor associated with a caffeine solution 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5
0.045% citric acid solution = 3.0
0.065% citric acid solution = 5.0
Buttery Sweet, dairy-like aromatic associated with natural butter Land O’Lakes unsalted butter tasted = 7.0 (F and A)
Earthy Odor of humid dirt or potting soil (A and F)
Boiled vegetable Flavor like cooked vegetable
Seafood Flavor like seafood
Smokiness Intensity of smoke-like (A & F)
Feeling factors
Astringent The sensation on the tongue, described as puckering/dry, and associated with strong tea
Metallic The sensation on the tongue described as flat and associated with iron and copper
Peppery The sensation on the tongue described as tingly and associated with pepper
Texture
Juiciness The amount of perceived juice that is released from the product during mastication Carrot = 8.5
Mushroom = 10.0
Cucumber = 13.5
Apple = 13.5
Watermelon = 15.0
Beef strip loin grilled to 58°C = 9.0
Beef strip loin grilled to 80°C = 4.0
Chewiness Length of time required to masticate product (at a constant rate of force) to reduce it to a consistency suitable for swallowing Marshmallows = 3.0
Nibs Twizzlers = 5.0
Tootsie Roll = 9.0
Starburst = 12.0
Haribo gummy bears = 14
Firmness Force required to compress sample between tongue and palate Cream cheese = 1.0
Egg white = 2.5
Yellow American cheese = 4.5
Olives = 6.0
Frankfurter = 7.0
Pastiness Degree in which fish turns in to a paste after chewing
Springiness Springiness of salmon on first few bites
Light texture Lightness and smoothness of salmon in mouth
Toothpacking or tooth adherence The degree to which the fish sticks between molars after mastication Carrots = 1.0
Mushrooms = 3.0
Saltine crackers = 5.0
Graham cracker = 7.5
Cheese puffs = 11.0
Ju-Jubes candy = 15.0
Clinginess Degree to which the salmon clings to the mouth and teeth
Flakiness Degree to which the salmon breaks down into flakes in the mouth
Crumbliness Degree of fish disintegration in the first bite
Tenderness Tenderness of salmon on first few bites
Aftertaste
Time Time when aftertaste starts
Overall aftertaste Intensity of aftertaste
Earthy aftertaste Intensity of earthy aftertaste
Metallic aftertaste Intensity of metallic aftertaste
Chicken-like aftertaste Intensity of chicken-like aftertaste
Oily aftertaste Intensity of fish oil aftertaste
  • Abbreviations: A, aroma; F, flavor; TMA, XXX.

Table 8.

Definition and reference standards for meat-descriptive flavor, basic tastes, and texture attributes for ground beef, adapted from Adhikari et al. (2011)

Attributes Definition Reference
Apricot Fruity aromatics that can be described as specifically apricot Sun sweet dried apricot = 7.5 (F)
Asparagus The slightly brown, slightly earthy green aromatics associated with cooked green asparagus Asparagus water = 6.5 (F), 7.5 (A)
Animal hair The aromatics perceived when raw wool is saturate with water Caproic acid = 12.0
Barnyard Combination of pungent, slightly sour, hay-like aromatics, associated with farm animals and inside of a horn White pepper in water = 4.0 (F), 4.5 (A)
Tinure of civet = 6.0 (A)
Beef identity Amount of beef-flavor identity in the sample Swanson’s beef broth = 5.0
80% lean ground beef = 7.0
Beef brisket = 11.0
Beet A dark damp-musty-earthy note associated Food Club sliced beets juice with 1-part juice with canned red beets to 2 parts water = 4.0 (F)
Bitter The fundamental taste factor associated with a caffeine solution 0.01% caffeine solution = 2.0
0.02% caffeine solution = 3.5
Bloody/serum-like The aromatics associated with blood on cooked meat products; closely related to metallic aromatic USDA Choice strip steak = 5.5
Beef brisket = 6.0
Brown/roasted A round, full aromatic generally associated with beef suet that has been broiled Beef suet = 8.0
80% lean ground beef = 10.0
Buttery Sweet, dairy-like aromatic associated with natural butter Land O’Lakes unsalted butter = 7.0
Burnt The sharp/acrid flavor note associated with over-roasted beef muscle, something over-baked or excessively browned in oil Alf’s red wheat Puffs = 5.0
Chemical The aromatics associated with garden hose, hot Teflon pan, plastic packaging, and petroleum-based product such as charcoal Ziploc sandwich bag = 13.0
Clorox in water = 6.5
Chocolate/cocoa The aromatics associated with cocoa beans and powdered cocoa and chocolate bars; brown, sweet, dusty, often bitter aromatics Hershey’s cocoa powder in water = 3.0
Hershey’s chocolate kiss = 8.5
Cooked milk A combination of sweet, brown flavor notes and aromatics associated with heated milk Mini Babybel original Swiss cheese = 2.5
Dillon’s whole milk = 4.5
Cumin The aromatics commonly associated with cumin and characterized as dry, pungent, and woody; slightly floral McCormick ground cumin = 7.0 (F), 10.0 (A)
Diary The aromatics associated with products made from cow’s milk, such as cream, milk, sour cream or butter milk Dillon’s reduced fat milk (2%) = 8.0
Fat-like The aromatics associated with cooked animal fat Hillshire Farms Lit’l Beef Smokies = 7.0
Beef suet = 12.0
Floral Sweet light, slightly perfume impression associated with flowers Welch’s white grape juice in water = 5.0
Geraniol = 7.5 (A)
Green Sharp, slightly pungent aromatics associated with green/plant/vegetable matters such as parsley, spinach, pea pod, fresh-cut grass, etc. Hexanal in propylene glycol (5000 ppm) = 6.5 (A)
Fresh parsley water = 9.0
Green-hay-like Brown/green dusty aromatics associated with dry grasses, hay, dry parsley, and tea leaves Dry parsley in medium snifter = 5.0 (A)
Dry parsley in ∼30-mL cup = 6.0
Heated Oil The aromatics associated with oil heated to a high temperature Wesson oil, microwaved 3 min = 7.0
Lay’s potato chips = 4.0 (A)
Leather Musty, old leather (like old book bindings) 2,3,4-trimethoxybenzaldehyde = 3.0(A)
Liver-like The aromatics associated with cooked organ meat/liver Beef liver = 7.5
Oscar Mayer Braunschweiger
Medicinal A clean sterile aromatic characteristic of antiseptic-like products such as Band-Aids, alcohol, and iodine Band-Aid = 6.0 (A)
Metallic The impression of slightly oxidized metal, such as iron, copper, and silver spoons 0.10% potassium chloride solution = 1.5
USDA Choice strip steak = 4.0
Dole canned pineapple juice = 6.0
Musty/earthy/humus Musty, sweet, decaying vegetation Sliced button mushrooms = 3.0 (F and A)
1000 ppm of 2,6-dimethcycyclohexanol in propylene glycol = 9.0 (A)
Overall sweet A combination of sweet taste and sweet aromatics; the aromatics associated with the impression of sweet Post shredded wheat, spoon size =1.5 (F)
Hillshire Farms Lit’l Beef Smokies = 3.0
Petroleum-like A specific chemical aromatic associated with crude oil and its refined products that have heavy oil characteristics Vaseline petroleum jelly = 3.0 (A)
Rancid The aromatics commonly associated with oxidized fat and oils; these aromatics may include cardboard, paint-like, varnish, and fishy Microwaved Wesson vegetable oil (3 min) = 7.0
Microwaved Wesson vegetable oil (5 min) = 9.0
Refrigerator/stale Aromatics associated with products left in refrigerator for an extended period of time and absorbing a combination of odors (lack of freshness/flat) 80% lean ground beef, stored overnight, and served at room temperature = 4.5 (F), 5.5 (A)
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Smoky charcoal An aromatic associated with meat juices and fat drippings on hot coats, which can be acrid, sour, burnt, etc. Wright’s Natural Hickory seasonings in water = 9.0 (A)
Smoky wood Dry, dusty aromatic reminiscent of burning wood Wright’s Natural Hickory seasoning in water = 7.5 (A)
Soapy An aromatic commonly found in unscented hand soap Ivory bar soap in water = 6.5 (A)
Sour aromatics The aromatics associated with sour substances Dillon’s buttermilk = 5.0
Sour milk/sour dairy Sour, fermented aromatics associated with dairy products such as buttermilk and sour cream Laughing cow light Swiss cheese = 7.0
Dillon’s buttermilk = 9.0
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5
0.050% citric acid solution = 3.5
Spoiled-putrid The presence of inappropriate aromatics and flavors that is commonly associated with the products; it is a foul taste and/or smell that indicates the product is starting to decay and putrefy Dimethyl disulfide in propylene glycol 10,000 ppm) = 12.0 (aroma)
Sweet The fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0
Umami Flat, salty, somewhat brothy; the taste of glutamate, salts of amino acids, and other molecules called nucleotides 0.035% accent flavor enhancer solution = 7.5
Warmed-over Perception of a product that has been previously cooked and reheated 80% lean ground beef (reheated) = 6.0
Cohesiveness of mass The degree to which chewed sample holds (at 10–15 chews) together in a mass Licorice = 0.0
Carrots = 2.0
Mushrooms = 4.0
Frankfurter = 7.0
American process cheese = 9.0
Soft brownie = 13.0
Pillsbury biscuit dough = 15.0
Mouthcoating A sensation of having a slick/fatty coating on the tongue and other mouth surfaces Half and half = 4.5
Whipping cream = 8.0
Hardness The force to attain a given deformation, such as the following: force to compress with the molars, force to compress between tongue and palate, force to bite through with incisors Cream cheese = 1.0
Egg white = 2.5
Yellow American cheese = 4.5
Olives = 6.0
Hebrew National frankfurter = 7.0
Planters peanut = 9.5
Life Savers = 14.5
Initial juiciness The amount of perceived juice that is released from the product during the initial 2–3 chews Carrot = 8.5
Mushroom = 10.0
Cucumber = 12.0
Apple = 13.5
Watermelon = 15.0
Particle size The degree to how large or small the particle is Small pearly tapioca = 4.0
Boba tea tapioca = 8.0
Springiness The degree to which sample returns to original shape or the rate with which sample returns to original shape Cream cheese = 0.0
Frankfurter = 5.0
Marshmallow = 9.5
Gelatin dessert = 15.0
  • Abbreviations: A, aroma; F, flavor.

Table 9.

Definitions and references for frankfurter and sausage-type product descriptive attributes and scaling intensities adapted from Knight (2006) and Nuñez De Gonzalez et al. (2004)

Attribute Definition References
Aromatics
Overall meat flavor Aromatics associated with meat flavor in general; can be any meat species Swanson® beef broth = 5.0 (F)
80% lean ground beef = 7.0 (A and F)
Beef brisket = 11.0 (A and F)
Fat-like Aromatics associated with cooked animal fat Pork fat, cooked = 10.0 (F), 7.0 (A)
Smoke Dry, dusty aromatic reminiscent of burning wood Wright’s Natural Hickory seasoning in water = 7.5 (A)
Spice Aromatics associated with the spices blend added to the product Place a small amount of spice blend on tongue = 15
Cardboard Aromatic associated with slightly oxidized fats and oils, reminiscent of wet-cardboard packaging Dry cardboard = 5.0 (F), 3.0 (A)
Wet cardboard = 7.0 (F), 6.0 (A)
Paint-like Aromatic associated with extensively oxidized fats and oils, paint thinner, varnish Old beef tallow = 10
Old Pringles = 13
Fishy Aromatics associated with fish Canned StarKist tuna = 12 (F), 10 (A)
Canned chicken = 4 (F)
Caramelized A distinct sweet and rich flavor, with notes of toffee, butterscotch, and molasses Werther’s hard caramel = 15 (F)
Vinegary Aroma notes associated with vinegar 1.1g vinegar in 200 g water = 6.0 (F), 4.0 (A)
Feeling factors
Astringent The chemical feeling factor on the tongue or other skin surfaces of the oral cavity, described as puckering/dry and associated with tannins or alum Lipton tea, 1 bag = 6.0 (F)
Lipton tea, 3 bags = 12.0 (F)
Metallic The impression of slightly oxidized metal, such as iron, copper, and silver spoons Dole pineapple juice = 6.0 (A and F)
0.10% KCl solution = 1.5 (A and F)
Basic tastes
Salt Aromatics associated with products left in the refrigerator for an extended period time and absorbing a combination of odors (lack of freshness/flat) 80/20 ground pork, cooked, left chilled overnight = 6.0 (F), 8.0 (A)
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5
0.045% citric acid solution = 3.0
0.065% citric acid solution = 5.0
Bitter The fundamental taste factor, associated with a caffeine solution 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sweet The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Aftertastes
Fat mouthfeel Aromatics associated with cooked animal fat after expectoration of the sample Pork fat, cooked = 10.0 (F), 7.0 (A)
Sour The fundamental taste factor associated with citric acid that remains in the mouth after expectoration 0.015% citric acid solution= 1.5
0.045% citric acid solution = 3.0
0.065% citric acid solution = 5.0
Bitter The fundamental taste factor associated with a caffeine solution that remains in the mouth after expectoration 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Salty The fundamental taste factor of which sodium chloride is typical and that remains in the mouth after expectoration 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Textures
Springiness The degree to which the sample returns to original shape or the rate with which the sample returns to original shape Cream cheese = 0.0
Frankfurter = 5.0
Marshmallow = 9.5
Gelatin dessert = 15.0
Juiciness The amount of perceived juice that is released from the product during mastication Carrot = 8.5
Mushroom = 10.0
Cucumber = 13.5
Apple = 13.5
Watermelon = 15.0
Hardness The force to attain a given deformation, such as the following: force to compress with the molars, force to compress between tongue and palate, force to bite through with incisors Cream cheese = 1.0
Egg white = 2.5
Yellow American cheese = 4.5
Olives = 6.0
Hebrew National frankfurter = 7.0
Planters peanut = 9.5
Life Savers = 14.5
Cohesiveness of mass The degree to which chewed sample holds (at 10–15 chews) together in a mass Licorice = 0.0
Carrots = 2.0
Mushrooms = 4.0
Frankfurter = 7.0
American process cheese = 9.0
Soft brownie = 13.0
Pillsbury biscuit dough = 15.0
  • Abbreviations: A, aroma; F, flavor.

Table 10.

Definition and reference standards for descriptive attributes for pork crumbles or other processed meat adapted from Waters (2017)

Attribute Definition References
BHA/BHT The aromatics associated with the chemicals butylated hydroxytoluene 0.01% BHA/0.01% BHT = 12.0, hydroxyanisole and butylated (A)
Browned/roasted A round, full aromatic generally associated with pork suet that has been broiled Pork suet = 8.0 (taste)
Fresh ground pork = 10.0 (taste)
Burnt The sharp/acrid flavor note, associated with over-roasted beef muscle, something over-baked or excessively browned in oil Alf’s red wheat puffs = 5.0 (taste)
Buttery Sweet, dairy-like aromatics associated with natural butter Land O’Lakes unsalted butter = 7.0 (A and F)
Cardboardy The fundamental taste factor associated with cardboard Cardboard (wet) = 6.0 (A and F)
Cardboard (dry) = 4.0 (A), 5.0 (F)
Chemical The aromatic associated with garden hose, hot Teflon pan, plastic packaging, and petroleum-based products Ziploc sandwich bag = 13.0 (A)
Clorox in water = 6.5 (F)
Fat-like The aromatic associated with cooked animal fat Hillshire Farms Lil’ Beef Smokies = 7.0 (F)
Pork suet = 12.0 (taste)
Green-hay-like Sharp, slightly pungent aromatics associated with green/plant/vegetable matter such as parsley, spinach, pea pod, fresh-cut grass, etc. Hexanal in propylene glycol (5000 ppm) = 6.5 (A)
Fresh parsley water = 9.0 (F)
Heated oil The aromatics associated with oil heated to a high temperature Lay’s potato chips = 4.0 (A)
Wesson vegetable oil = 7.0 (F)
Metallic The impression of slightly oxidized metal, such as iron, copper, silver spoons placed in mouth 0.10% potassium chloride solution = 1.5 (F)
Overall sweet The combination of sweet taste and sweet aromatics Post shredded wheat = 1.5 (F)
Hillshire Farms Lil’ Beef Smokies = 3.5 (F)
Petroleum-like A specific chemical aromatic associated with crude oil and its refined products that have heavy oil characteristics Vaseline petroleum jelly = 3.0 (A)
Pork flavor Amount of pork-flavor identity in the sample Fresh ground pork = 8.0 (F)
Rancid An aromatic commonly associated with oxidized fat and oils; these aromatics may include cardboard, paint-like, varnish, and fishy Wesson vegetable oil (3 min) = 7.0 (F)
Wesson vegetable oil (5 min) = 9.0 (F)
Refrigerator stale Aromatics associated with products left in refrigerator for an extended period of time and absorbing a combination of odors (lack of freshness/stale) Ground pork (1-d old) = 5.5 (A), 4.5 (F)
Rosemary The aromatics associated with rosemary extract 0.02% rosemary extract = 12.0 (A)
Smoky wood Dry, dusty aromatic reminiscent of burning wood Wright’s Natural Hickory seasoning in water = 7.5 (A)
Sorghum The fundamental aromatic and taste factor associated with sorghum bran Ground pork with sorghum bran added = 7.0 (F)
Sour aromatics The aromatics associated with a sucrose solution Dillon’s buttermilk = 5.0 (A)
Sour milk/dairy Sour, fermented aromatics associated with dairy products such as buttermilk and sour cream Laughing Cow light Swiss cheese = 3.0 (A), 7.0 (F)
Dillon’s buttermilk = 4.0 (A), 9.0 (F)
Spice complex The fundamental taste factor from a specific spice blend Spice complex = 12.0 (F)
Spoiled The presence of inappropriate aromatics and flavors that is commonly associated with the products; it is a foul taste and/or smell that indicates the product is starting to decay and putrefy Dimethyl disulfide in propylene glycol (10,000 ppm) = 12.0 (A)
Warmed-over Perception of a product that has been previously cooked and reheated Ground pork = 6.0 (A and F)
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Salty The fundamental taste factor of which sodium chloride is typical 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Sweet The fundamental taste factor associated with sucrose 2.0% sucrose solution = 2.0
Bitter The fundamental taste factor associated with a caffeine solution 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sour The fundamental taste factor associated with citric acid 0.015% citric acid solution = 1.5
0.045% citric acid solution = 3.0
0.065% citric acid solution = 5.0
Umami Flat, salty, somewhat brothy; the taste of glutamate, salts of amino acids and other molecules called nucleotides 0.035% accent flavor enhancer solution = 7.5
  • Abbreviations: A, aroma; BHA, butylated hydroxyanisole; BHT, butylated hydroxytoluene; F, flavor.

Table 11.

Definitions and references for cured bacon descriptive attributes and scaling intensities adapted from Jeremiah et al. (1996) and Meilgaard et al. (2025)

Attribute Description References
Aromatics
Meaty/brothy complex The brown aromas or flavors associated with broth but not with any particular species
Bacon identity The combined flavor reminiscent of cooked bacon
Browned A round, full aromatic generally associated with pork suet that has been broiled Pork suet = 8.0 (taste)
Fresh ground pork = 10.0 (taste)
Fat complex The aromas and flavor associated with meat fat but not with any particular species Hillshire Farms Lil’ Beef Smokies = 7.0 (F)
Pork suet = 12.0 (taste)
Smoke The dark brown aromas and flavors associated with burning or charred wood Wright’s Natural Hickory seasoning in water = 7.5 (A)
Maple Flavor reminiscent of maple extract, maple syrup Mrs. Butterworth’s original syrup = 10
Spice complex The aromas and flavors associated with spices but not any particular spice Place a pinch of spice blend used in the product on the tongue = 15
Feeling factors
Metallic The sensations on the tongue associated with metals such as iron or copper 0.10% potassium chloride solution = 1.5 (F)
Silver spoons placed in mouth
Astringent The chemical feeling factor on the tongue or other skin surfaces of the oral cavity, described as a puckering/dry and associated with tannins or alum Lipton tea, 1 bag = 6.0 (F)
Lipton tea, 3 bags = 12.0 (F)
Mouthburn The burning sensation of the mouth caused by substances such as capsaicin or piperdine
Tastes
Salt The taste stimulated by sodium salts such as sodium chloride 0.15% sodium chloride solution = 1.5
0.25% sodium chloride solution = 3.5
Sour The taste stimulated by acids such as citric or malic acid 0.015% citric acid solution= 1.5
0.045% citric acid solution= 3.0
0.065% citric acid solution = 5.0
Bitter The taste stimulated by such substances such as quinine or caffeine 0.05% caffeine in 1000 mL water = 2.0 (F)
0.08% caffeine in 1000 mL water = 5.0 (F)
Sweet The taste stimulated by sugars such as sucrose or fructose 2.0% sucrose solution = 2.0
Aftertastes
Fat mouthfeel The amount of oily film remaining on the oral palate after expectorating Whipping cream = 8.0
Cream cheese = 1.0
Pork fat, cooked = 10.0(F), 7.0(A)
Smoke/spice complex The amount of smoke/spice complex aromas or flavors remaining after expectorating The remaining aromatic in the mouth after expectoration of the sample
Texture
Springiness The amount a sample returns to original shape after a certain time period Cream cheese = 0.0
Frankfurter = 5.0
Marshmallow = 9.5
Gelatin dessert = 15.0
Hardness The force required to bite through a sample Cream cheese = 1.0
Egg white = 2.5
Yellow American cheese = 4.5
Olives = 6.0
Hebrew National frankfurter = 7.0
Planters peanut = 9.5
Life Savers = 14.5
Juiciness The amount of wetness released from the sample upon biting Carrot = 8.5
Mushroom = 10.0
Cucumber = 12.0
Apple = 13.5
Watermelon = 15.0
Cohesiveness of mass The degree to which the sample mass holds together after chewing Licorice = 0.0
Carrots = 2.0
Mushrooms = 4.0
Frankfurter = 7.0
American process cheese = 9.0
Soft brownie = 13.0
Pillsbury biscuit dough = 15.0
  • Abbreviations: A, aroma; F, flavor.

Table 12.

List of potential descriptive sensory attributes for processed, breaded chicken patties, and nuggets adapted from Lyon et al. (1980) and Baublits et al. (2006)

Visual Appearance Attributes of the Breading Flavor Aromatics of the Meat Block Texture With 1st Bite
Color Chicken identity Hardness
Surface roughness/smoothness Fat-like Springiness
Uniformity of breading surface Wet feathers Cohesiveness
Cardboardy Chewiness
Odor Soapy Juiciness
Chicken identity Fishy Crispness
Browned Paint-like Fracturability
Heated oil Chemical
Wet feathers Chew down characteristics
Cardboardy Basic tastes Cohesiveness of mass
Paint-like Salt Hardness of mass
Fishy Sour Juiciness of mass
Bitter Number of chews
Flavor aromatics of the breading Sweet
Browned Residuals
Fat-like/Oiliness Aftertastes Oily/greasy film
Heated oil Fat mouthfeel Residuals
Flour Sour Particle size
Soapy Bitter
Fishy Salty Others: for any class of attributes
Paint-like
Chemical
Table 13.

Definitions and references for smoked food products descriptive attributes and scaling intensities adapted from Jeffe et al. (2017)

Attribute Definition and References With Preparation References
Smoky overall An aromatic that may present characteristics of sweet, brown, pungent, acrid, slightly ashy or charred/burnt Tones Liquid Smoke (2 drops in 300 mL water) = 4.0 (A); mix 2 drops of Tones liquid smoke with 300 mL water, serve in a large 24-oz snifter, cover
Tones liquid smoke (2 drops in 200 mL water) = 5.5 (A); mix 2 drops of Tones liquid smoke with 200 mL water, serve in a large 24-oz snifter, cover
Tones liquid smoke (2 drops in 100 mL water) = 7.0 (A); mix 2 drops of Tones liquid smoke with 100 mL water, serve in a large 24-oz snifter, cover
Wood ashes = 5.0 (A) (wood ashes vary in their intensity depending on the type of wood and degree of burn but can be used if the liquid smoke reference is not available); obtain ashes from burnt wood (from fireplace or outdoor fire pit); place ashes in 2-oz glass jars with screw-on type lids; fill jars approximately 1/3 full; this may be prepared several days in advance and stored at room temperature, tightly sealed; prepare 1 jar for every 3 participants; these will be shared for smelling only
Ashy Dry, dusty, dirty smoky aromatics associated with the residual of burnt products Gerkens Midnight Black (BL80) cocoa powder = 3.5(F); mix 1 = 4 tsp of cocoa powder with 100 mL of water; serve in 1-oz cup
Woody The sweet, brown, musty, dark aromatics associated with the bark of a tree Diamond shelled walnuts = 4.0 (F); serve walnuts in a 1-oz cup
Musty/dusty The aromatics associated with dry, closed air spaces such as attics and closets; may be dry, musty, papery, dry soil or grain Kretschner wheat germ = 5.0 (A and F); aroma: serve 1 tablespoon wheat germ in a medium 12-oz snifter, cover; flavor: serve wheat germ in a 1-oz cup
Musty/earthy Somewhat sweet, heavy aromatics associated with decaying vegetation and damp black soil 50,000 ppm 1, 2, 4 trimethoxybenzene = 4.0 (A)
4000 ppm geosmin = 9 (A); for the first 2 chemicals, dilute the chemicals in deionized water and then dip a perfumer’s strip into the solutions and place in capped test tubes; place 1 drop of essence on a cotton ball in a large 24-oz snifter; cover
Miracle-Gro potting mix soil = 9.0 (A); fill a 2-oz glass jar half full with potting soil and seal tightly with screw-on type lid
Le Nez du Café No. 1 “earthy” = 12.0 (A)
Burnt The dark brown carbon impression of an over-cooked or over-roasted product that can be sharp, bitter, and sour Mixture of raw and burnt peanuts (3 raw:1 burnt) = 7.5 (A)
Mixture of raw and burnt peanuts (1 raw:1 burnt) = 10.0 (A)
Burnt peanuts = 13.0 (A)
For burnt peanuts: preheat oven to 218°C. Place 1 c of blanched peanuts in 3-L glass pan in a single layer; roast for 30 min—stir roasted peanuts every 5 min; let roasted (burnt) peanut completely cool; Blend the ratio of raw and burnt peanuts specified above in a food processor until small particles are formed; serve 15 mL in a large 24-oz snifter
Acrid The sharp, pungent, bitter, acidic aromatics associated with products that are excessively roasted or browned Wright’s Liquid Smoke Mesquite (1 drop in 100 mL water) = 3.0 (A); mix 1 drop of Wright’s Liquid Smoke Mesquite with 100 mL water, serve in a large 24-oz snifter, cover
Wright’s Liquid Smoke Mesquite (1 drop on cotton ball) = 9.5 (A); place 1 drop of liquid smoke on a cotton ball in a large 24-oz snifter; cover
Pungent The sharp physically penetrating sensation in the nasal cavity Reese horseradish and Hiland sour cream = 2.5 (A); mix 1 g of sour cream with 0.68 g of horseradish; serve in a medium snifter; cover
Heinz white vinegar = 8.0 (F); mix 1 part vinegar with 8 parts water; serve in a 1-oz cup
S&B wasabi paste in water = 10.0 (A); mix 1 g of wasabi paste in 50 mL water; serve in a medium 12-oz snifter; cover
Petroleum-like A specific chemical aromatic associated with crude oil and its refined products that have heavy oil characteristics Vaseline petroleum jelly: 3.0 (A); place a tsp of jelly in a medium 12-oz covered snifter
Motor oil, Castrol HD Optimum viscosity = 7.5 (A); place 1 drop of motor oil on a cotton ball in a medium 12-oz snifter; cover
Creosote/tar A pungent chemical aromatic associated with unrefined crude oil products Tar = 12.0; put 1 tsp of tar into a 60 mL glass jar and seal tightly with a screw-on type lid
Cedar A slightly sweet, dark, and musty aromatic associated with trees such as cedar Hartz Natural Red Cedar Small Animal Bedding & Litter = 6.0 (A); place 5 g of red cedar shaving in a medium 12-oz snifter; cover
Aldrich cedarwood oil, Virginia = 9.5 (A); place 1 drop of cedarwood oil on a cotton ball in a large 24-oz snifter; cover
Bitter The fundamental taste factor associated with a caffeine solution 0.01% caffeine solution = 2.0
0.02% caffeine solution = 3.5
0.035% caffeine solution = 5.0
0.05% caffeine solution = 6.5
Metallic An aromatic and mouthfeel associated with tin cans or aluminum foil 0.10% potassium chloride solution = 1.5 (F)
Sour The fundamental taste factor associated with a citric acid solution 0.015% citric acid solution = 1.5
0.05% citric acid solution = 3.5
  • Abbreviations: A, aroma; F, flavor.

These are living lexicons and are only examples of attributes reported. If panelists are trained to detect attributes on the lexicon, attributes may not be detected or new attributes may be defined during training for a specific experiment. “Other” should always be defined on a ballot so that attributes that may not have been detected in training but are apparent in testing can be identified. Attributes used in training can be included on the ballot so that panelists consider their presence as well.

Because processed meats are variable in sensory attributes across product type, a full complement of lexicons is difficult to present. However, examples of sensory attributes for ground beef, frankfurters, pork sausage crumbles, bacon, and chicken nuggets are presented in Tables 8, 9, 10, 11, and 12, respectively. A lexicon to describe the smoky character in smoked food products is presented in Table 13. For processed meats, sensory attributes can be derived from ballot development sessions, expanded from the whole-muscle lexicons, and based on previously reported attributes defined in scientific literature. It is important that each attribute be defined and references provided to assist in identifying and scaling the attribute. As many processed meat products contain added ingredients, the addition of aroma, flavor, basic taste, texture, and/or aftertastes associated with the addition of nonmeat ingredients should be included in the training, ballot development sessions, and sensory testing. Seasonings may be defined based on individual components that are contained in the seasoning (i.e., garlic flavor, pepper flavor, etc.) or as a combined attribute of spice complex. For example, in Table 9 the attributes of butylated hydroxyanisole/butylated hydroxytoluene, rosemary, and sorghum were added since they represent added ingredients in the study.

All ballots should include an attribute for “other.” During ballot development sessions and training, an attribute may not be expressed. By including an attribute “other,” panelists may identify attributes that are present during testing but that were not detected during training. The panel leader should monitor the other attribute category daily, discuss with panelists what they are perceiving, provide potential references, and add a new attribute to the ballot if appropriate.

Ballot development sessions are where trained panelists who have experience in attribute development determine attributes within a study. Treatments or products that will be included in the study and similar products to provide a wide range of potential attributes should be included in these sessions (Meilgaard et al., 2016). Note that attributes that are volatile, or detected by the olfactory bulb, can be used as either odor attributes as identified through the nose, flavor aromatics as sensed when placed or chewed in the mouth, or as aftertastes that are detected after either swallowing or expectorating the sample. Attributes can include visual, odor, flavor aromatics, basic tastes, texture, and aftertaste attributes.

Selection and training sensory panelists for discriminative or descriptive testing

Before initiating panelist training, the researcher must determine which sensory test method is most appropriate to fulfill the objectives of the study. Various test methods are available, and it is not the intent of these guidelines to fully describe these procedures. It is suggested that ASTM MNL26-3RD-EB (2020) and Meilgaard et al. (2025) be reviewed to determine the most appropriate procedure. A short discussion of the most common methods used was presented in the previous sections. Panelists can be trained to participate in discriminative and descriptive testing.

Selection of potential panelists

Trained panelists can be recruited from the surrounding community (external panelists), or they can be company employees (internal panelists). There are advantages and disadvantages to each type of panelist.

The advantage of using externally trained sensory panelists is that they are independent and do not have institutional or product knowledge or product loyalty. Product-specific attributes that could induce expectation errors with internal panelists are easily avoided. Additionally, a large selection pool is available when using external panelists. Because external panelists have only one obligation within the organization, they may have fewer daily conflicts. The main disadvantage of using external panelists is that they are not on site. Parking and ease of travel to the testing facility have to be considered.

Internal panelists are easily recruited within a company or research entity, and the prospect of excused time from daily responsibilities can be a motivating factor. The time requirement, however, needs to be communicated effectively to potential panelists and supervisors and supported by upper management. The selection pool for qualified candidates might be limited, and the tendency to accept marginal panelists to increase the number of panelists should be avoided. A key advantage of using internal panelists, such as Research & Development and Quality Assurance/Quality Control personnel and/or graduate students, is that their scientific knowledge and understanding makes them easier to train and they tend to have the ability to make more concise judgments. Furthermore, training in sensory evaluation enhances their ability to detect and describe flavor and/or texture attributes and makes them more conscious of sensory attributes in Research & Development and Quality Assurance/Control.

Recruitment of internal and external candidates may occur through advertisements, social media, posted announcements, and formal and informal verbal methods. When written recruitment methods are used, there are frequent word limitations, thus it is critical that the important factors are clearly identified. For external candidates, the advertisement should plainly indicate the time commitment required and if it is a paid position. The wording should focus on seeking those who have an interest in food, clearly state the importance of the panel’s effort, explain any limitations (food allergies or religious concerns), and include that training is a component.

Prescreening candidates

Potential candidates are first prescreened to determine their level of interest, availability, dependability, health (including dentures, allergies, use of medication), work experience, gender, age, smoking/tobacco use status, and food likes/dislikes (Figure 25). Background information on prospective panelists is valuable in selecting those individuals who have the greatest potential to become effective panelists.

Figure 25.
Figure 25.
Figure 25.
Figure 25.

Example of prescreening questionnaire for selection of panelists for a meat texture and flavor descriptive attribute sensory panel (modified from Meilgaard et al., 2016).

During the prescreening process, a candidate’s ability to follow directions or make concise judgments should also be determined because panelists who are not able to do these tasks will not be successful panelists. Logic tests can be conducted to determine a candidate’s ability to follow directions and make decisions. Figure 26 provides an example of logic tests, and additional logic tests can be found in Meilgaard et al. (2016). Candidates should mark the line in the approximate area for a correct response. Those who provide incorrect answers by marking the line in the opposite direction as requested should be immediately eliminated. These candidates will not follow instructions, and the panel leader will have to continually work with them to ensure that they understand. The result can be a decrease in motivation and positive attitude of other panelists. When grading, a 70% or above is considered acceptable on a logic test.

Figure 26.
Figure 26.

Examples of logic tests used for screening potential sensory panelists. Adapted from Meilgaard et al. (2016).

Screening candidates

Candidates who have passed the prescreening stage are invited to participate in a series of screening exercises. The purpose of screening is to identify candidates who have the required sensory and cognitive skills to perform the descriptive analysis tasks (ASTM MNL90-2ND-EB, 2023). Candidates must therefore be screened for the ability to (1) detect differences in product characteristics and in their intensities, (2) describe those characteristics using verbal descriptors and scaling methods for the different intensity levels, and (3) perform abstract reasoning as descriptive analysis relies heavily on the use of references when characteristics must be quickly recalled and applied to other products (Meilgaard et al., 2025).

It is often necessary to screen 4 times more panelists than what will be needed in the actual panel. Sensory protocols and procedures during screening should be similar to those used later in actual studies. It is important to use the same meat products for screening that will be evaluated by the panel after training. A higher level of discrimination is required for descriptive analysis panelists than for discriminative tests. The screening procedures should be more rigorous for descriptive analysis panelists and will require that more panelists be screened to achieve the final number of panelists needed. Screening panelists can include more than one of the strategies discussed below.

Matching tests are designed to determine a potential panelist’s ability to describe or identify descriptive attributes. Meat products contain multiple flavors and textures, especially further-processed products. Even if the original objective of the trained panel is to evaluate whole-muscle beef steaks, panelists need to have the ability to describe, identify, and rate flavor attributes. If the panelists can describe texture attributes of a further-processed meat product, such as a frankfurter or sausage product, the panel can be expanded or can be further trained to evaluate these products. Meilgaard et al. (2025) described these tests in detail. It is useful to present matching tests for taste and aroma to better assess a candidate’s ability to discriminate. Table 14 provides examples of matching tests from Meilgaard et al. (2025). The candidate is served a set of cups labeled sweet, sour, salty, bitter, and water, as well as a set of 5 cups coded with 3-digit random numbers. The labeled cups contain the appropriate stimulus at just more than threshold levels, and the coded cups contain a corresponding stimulus. The candidate is instructed to taste the labeled cups first to familiarize themselves with the basic taste. They are then asked to taste the coded samples and indicate on a score sheet the matching code number for each stimulus. Panelists should use palate cleansers between evaluating each sample.

Table 14.

Suggested samples for a matching taste test, adapted from Meilgaard et al. (2025)

Basic Taste Stimulus Concentration (g/L)
Sweet Sucrose 20
Salty Sodium chloride 2.0
Sour Citric acid 0.5
Bitter Caffeine 0.6
Astringent Alum 0.1%
Water Filtered water

Sniff tests are conducted for matching tests for aromas (Meilgaard et al., 2025). To conduct sniff tests, cut 0.75-cm-wide strips of Whatman filter paper at least 2.54-cm long. Cotton balls can be used as an alternative to Whatman filter paper. Obtain essential oils representing a variety of aromas from a business that uses essential oils to formulate food or personal care products (Table 15). Dip the filter paper strip in the oil sufficiently to concentrate the oil’s aroma. Allow the strip to dry under a hood for 30 min, then place it in a glass jar with a lid, a glass sniffer with a glass concave watch top, or in a test tube with a lid. Label the containers with random 3-digit codes. Ask the candidate to remove the lid or push the lid aside and sniff the contents of the container without touching the contents. Panelists can shake the container with the lid intact to concentrate the aromatics before evaluation. Candidates are asked to match the aroma they detect with the appropriate aroma descriptor selected from a list at the bottom of the score sheet; list 2 to 4 more descriptors than the number of stimuli presented. Candidates should be accepted if they are able to correctly identify approximately 70% of the attributes that are presented.

Table 15.

Suggested samples for sniff tests, adapted from Meilgaard et al. (2025)

Aroma Descriptors Stimulus
Anise, licorice Anise oil
Almond, cherry Amaretto, benzaldehyde, oil of bitter almond
Cinnamon Cinnamaldehyde, cassia oil
Clove, dentist’s office Eugenol, oil of clove
Ginger Ginger oil
Green, freshly mown lawn cis-3-Hexenol
Lemon, lime, or orange Lemon, lime, or orange oil
Peppermint, minty Peppermint oil
Vanilla Vanilla extract
Wintergreen Methyl salicylate, oil of wintergreen

Ranking tests are used to determine a candidate’s ability to discriminate graded levels of intensity of a given attribute. As described in Meilgaard et al. (2025), candidates are presented with a series of samples in random order. The samples are coded with 3-digit random numbers and cover a range of a specific attribute. Ask the candidates to rank the samples in order of increasing intensity of the stated attribute. Examples of sample sets are listed in Table 16.

Table 16.

Suggested sample sets for ranking tests, adapted from Meilgaard et al. (2025)

Attribute Sensory Stimuli Concentration (g/L)
Taste
 Sweet Sucrose/water 10 20 50 100
 Salty Sodium chloride 1 2 5 10
Texture
 Hardness Cream cheese,a American cheese,a peanuts, carrotb
 Juiciness Banana,b carrot,b mushroom,b appleb
  • 1.27-cm cubes.

  • 1.27-cm slices.

Sniff tests can also be used as identification tests to understand a panelist’s ability to recognize, describe, or identify descriptors. These tests are mainly used in the selection of descriptive panelists. Sniff tests, as described above, can be used. Jars with varying flavor aromatics are presented individually to panelists using 3-digit random codes. The candidate is asked to describe the attribute that they detect on a sheet of paper. When evaluating their score sheet, a correct answer is when the candidate uses descriptors similar to the attribute. For example, if cedar oil is used in the sniff test, correct answers could be cedar, wood, a forest, sweaters, stored sweaters, or a wooden chest.

Triangle tests are recommended to determine a panelist’s ability to discriminate differences in sensory attributes. Duo-Trio tests can be used for this exercise, but are not as sensitive, so more duo-trio tests are required to have the same sensitivity as triangle tests. For discriminative panel selection, fewer triangle tests need to be run (from 6 to 8); for descriptive panelists, a larger number of triangle tests should be conducted to provide a greater range of attributes (10 to 15).

A sequential analysis procedure is used to minimize the number of tests needed for screening (Bradley, 1953). That procedure facilitates an early decision on the suitability of panelists to participate as panel members. The sequential procedure makes one of the following decisions after each triangle test: 1) Accept the candidate as a potential panelist; 2) reject the candidate; or 3) continue testing. The decisions are based on specifications of 4 parameters: P0 = Maximum proportion of correct decisions ruled as an unacceptable candidate; P1 = Minimum proportion of correct decisions ruled as an acceptable candidate; α = Probability of selecting an unacceptable candidate; and β = Probability of rejecting an acceptable candidate. By plotting test numbers against the accumulated number of correct test results, a decision is made based on the region in which the point is plotted (Figure 27). The regional boundaries are described in greater detail by Cross et al. (1978) and Meilgaard et al. (2025).

Figure 27.
Figure 27.

Sequential analysis chart used for screening potential panelists. Adapted from Cross et al. (1978).

Examples of triangle tests used in the screening procedure are shown in Table 17. Test samples for triangle tests are prepared to give a 2-unit difference on an 8-point scale, or a 3- to 5-point difference on a 16-point scale for the attribute being tested. In this example, the attributes are tenderness, juiciness, and connective tissue amount. More than 5 triangle tests are recommended, but these 5 examples can be used as an example of how to create differences. Differences should be of a magnitude that could easily be detected by an experienced panel leader. The values P0 = 0.45, P1 = 0.70, α = 0.10, and β = 0.10 are used (Meilgaard et al., 2025). Figure 27 is then applied to accept, continue testing, or accept panelists.

Table 17.

Examples of triangle tests that can be used and the sensory attribute of interest

Triangle Tests Explanation of Attribute Differences Targeted for Testing
1. Strip loin steak cooked to 70°C and a strip loin steak cooked to 75°C, with juices pressed out to get differences in juiciness.
2. Eye of round steak, cooked to 70°C, and a strip loin steak cooked to 70°C, to get differences in tenderness.
3. Top sirloin steak and a strip loin steak, both cooked to 70°C to get differences in flavor intensity (or ST muscle “leached” in water for a few hours before cooking).
4. Ground beef patty and a ground beef patty containing 1.0% added ground liver, evenly mixed and distributed to get differences in liver flavor.
5. Brisket or bottom round steak grilled, and a strip loin steak grilled to get differences in connective tissue.
6. Chicken breast marinated with salt and phosphate, and chicken marinated with salt and a natural flavor phosphate replacer.
7. Pork sausage made from prerigor pork and postrigor pork.
  • Abbreviations: ST, Semitendinosus.

At the end of the screening period, the candidates in the “accept” region can be selected for training. If time is a factor, screening can be stopped after 10 sessions, but when evaluating for multiple attributes, a higher number of tests provides greater insight into the panelist’s ability to discriminate. It is desirable to select only those in the “accept” region, but if most of the candidates are in the “continue testing” region, they also could be selected. Another option would be to recruit more candidates and reinitiate the screening process. In panel selection, screening should not be considered a part of training, but rather a test to quickly eliminate those individuals who cannot detect large differences in attributes. At least twice as many individuals should be screened as are needed on the final panel.

Meilgaard et al. (2016) suggest that, when using triangle tests, to reject candidates scoring less than 60% correct on the easy tests or less than 40% on the moderately difficult tests. When using duo-trio tests, reject candidates scoring less than 75% on the easy tests or less than 60% on the moderately difficult tests. After completion of the screening tests, accept panelists for training who pass all stages (ASTM MNL90-2ND-EB, 2023). During candidate testing, observe panelist behavior, timeliness, ability to interact with others, confidence level, and dependability. Do not accept a panelist who misses an appointment, is difficult to work with, or is disruptive or domineering in a group setting. Panelists who are not dependable during a “job interview” will not be dependable after they have the job. If the panel leader determines that an individual is disruptive, the panelist’s behavior pattern will be difficult to alter and might take more effort than desired to change. Final panelists selected for training should be acceptable based on all criteria discussed above.

Training

The objectives of training are to familiarize an individual with test procedures, improve an individual’s ability to recognize and identify sensory attributes, and improve an individual’s sensitivity and memory, permitting precise and consistent sensory judgments. If using end-anchored or end and center-anchored line tests via a survey tool, a good training/practice is to also give panelists a number and have them mark the number on the line. They can practice this on their own.

Panelists should understand the importance of the study to ensure their cooperation and motivation. Let them know that you are pleased to have them participate and that their cooperation is appreciated. Without influencing the panelists’ future responses, give them as much specific information as possible on the purpose of the test. The importance of concentration should be stressed.

Panelists need to learn to be objective early in training. While all panelists are consumers, their opinions or preferences should not be expressed in their evaluations, nor should they influence others through discussion. Several decisions have to be reached early in training regarding protocols. The amount of sample that a judge places in his/her mouth must always be standardized. The decision of whether or not the panelist should swallow a sample should be standardized. To further standardize the evaluation methods, palate cleansing procedures should be established. After sample evaluation, taste bud refreshers and mouth rinsing should follow to minimize taste bud fatigue. Each panel member should rinse their mouth between samples. Room-temperature water—bottled, filtered, or distilled—is the most common rinse. When there is a great deal of aftertaste, taste bud refreshers such as unsalted crackers, plain bagels, seltzer water, apple slices (as long as a water rinse is used afterward to avoid flavor carryover), or ricotta cheese are useful. Taste bud refreshers should be used with caution. Even though they can eliminate lingering aromatics, mouthfeel, and aftertaste after evaluation of a product, taste bud refreshers can contribute to taste bud fatigue.

The interval between samples should be standardized and is dependent upon the product under study. Enough time should be allowed between samples to permit recovery from flavor buildup, yet not so much time that the panelist loses his/her ability to discriminate. Depending on the descriptive method used and level of experience with the product, the panel leader can provide panelists with predetermined lexicons as previously discussed, and with definitions and procedures for use during evaluation. Examples of training sessions for meat-descriptive palatability attributes and beef flavor are listed in Tables 18 and 19, respectively. In some instances, the panelists will develop their own descriptors and scaling techniques, such as with QDA. Training is best accomplished through individual and group sessions in which various samples of the product types that are representative of those in the sensory tests are evaluated and discussed. Multiple sessions should be devoted to demonstrating levels of each attribute under study, as shown in Tables 18 and 19. All attributes defined in Tables 18 and 19 may not need to be included in a study. The sensory professionals should determine the attributes to include and train for based on the objectives and treatments in the study.

Table 18.

Examples of training sessions for meat-descriptive palatability attributesa

Day Goals Exercises
1 1. To introduce the basic meat-descriptive attribute method to the panel Juiciness
Hand out the AMSA guidelines and a ballot. Explain what the meat-descriptive attribute sensory evaluation is.
2. To familiarize the panelists with the ballot Concentrate on understanding what juiciness is juicy vs. dry.
1. Strip loin steak—standard or warm-up sample; cooked to 70°C. Use this sample to ask panelists to just give an initial evaluation for juiciness with the panel leader scoring
3. To provide some initial descriptor for juiciness and MF tenderness 2. Strip loin steak—setting lower scale for juiciness; cooked to 80°C and pressed; evaluate juiciness with panel leader
3. Eye of round steak—cooked to 90°C; scaling between 1 and 2 for juiciness
4. Strip loin steak—standard or baseline steak; cooked to 70°C. Should get a similar score as 1 above
MF tenderness
Concentrate on understanding MF tenderness: tough vs. tender
5. Strip loin steak—standard or warm-up sample; cooked to 70°C; Use this sample to anchor on 5 or 6 of the scale
6. Tenderloin steaks—for a very tender sample; cooked to 65°C. Use this sample for a 7 or 8 on MF tenderness
7. Old cow steak—for a tough sample; cooked to 70°C. Use this to give a 2 or 3 for MF tenderness
2 1. To develop a baseline for juiciness and tenderness Juiciness
1. Strip loin steak—standard or warm-up sample; cooked to 70°C. Use this sample to ask panelists to just give an initial evaluation for juiciness with the panel leader scoring as an anchor
2. Strip loin steak—cooked to 75C and pressed; setting lower scale for juiciness
3. Eye of round steak—cooked to 85°C; scaling between 1 and 2 for juiciness
4. Strip loin steak—standard or baseline steak; cooked to 70°C
5. Strip loin steak—cooked to 60°C; scaling for higher juiciness
2. To begin scaling for each of the 2 attributes MF tenderness: tough vs. tender; also rate for juiciness to start combining
1. Strip loin steak—standard or warm-up sample; cooked to 70°C. Use this sample to anchor on 5 or 6 of the scale
2. Tenderloin steaks—for a very tender sample; cooked to 65°C. Use this sample for a 7 or 8 on MF tenderness
3. Old cow steak—for a tough sample; cooked to 70°C. Use this to give a 2 or 3 for MF
4. Eye of round steak—broiled to 85°C for the lower part of the scale
5. Strip loin steak—standard or warm-up sample; cooked to 70°C
3 1. To develop a baseline for juiciness and tenderness 1. Strip loin steak—standard or warm-up sample; cooked to 70°C. Use this sample to ask panelists to just give an initial evaluation for juiciness and MF tenderness
2. To begin scaling for each of the 2 attributes 2. Strip loin steak—cooked to 80°C and pressed; setting the lower scale for juiciness and toughness
3. To begin differentiating between MF tenderness and the amount of CT 3. Eye of round steak—cooked to 85°C, scaling between 1 and 2 for juiciness and toughness
4. To begin scaling for both attributes of tenderness 4. Tenderloin steaks—cooked to 65°C for a very tender sample
5. Old cow steak— cooked to 70°C for a tough sample
6. Strip loin steak—cooked to 60°C for higher juiciness
Introduce the concept of CT: need to separate tough vs. tender for MF tenderness and amount of CT
1. Strip loin steak—standard or warm-up sample, cooked to 70°C
2. Tenderloin steaks—cooked to 65°C for a very tender sample
3. Bottom round steak—to show CT
4. Old cow steak— cooked to 70°C for a tough sample
5. Broiled brisket steak—broiled to 75°C
6. Strip loin steak—standard or warm-up sample, cooked to 70°C
4 1. To begin, differentiate between MF tenderness and the amount of CT 1. Strip loin steak—standard or warm-up sample, cooked to 70°C
2. Cold-shortened steak—setting a lower scale for tenderness; cooked to 70°C. Describe MF tenderness. Describe the amount of CT
3. Top butt steak—scaling for tenderness, cooked to 70°C
2. To begin scaling for both attributes of tenderness 4. Tenderloin steak—scaling for tenderness, cooked to 70°C
5. Break
6. Strip loin steak—standard or warm-up sample, cooked to 70°C
7. Brisket steak—scaling for tenderness/CT, cooked to 70°C
8. Cold-shortened steak—scaling for tenderness, cooked to 70°C
9. Strip loin steak—scaling for tenderness and baseline determinations, cooked to 70°C
5 1. To fine-tune evaluation for juiciness, MF tenderness, and CT amount 1. Strip loin steak—standard or warm-up sample, cooked to 70°C
2. To scale for all attributes 2. Tenderloin steak—differences in flavor, CT, and juiciness
3. Top butt steak—differences in flavor, CT, and juiciness
4. Cold-shortened steak—differences in flavor, CT, tenderness, and juiciness
5. Strip loin steak—baseline evaluation
6. Water-soaked strip steak or eye of round—differences in all attributes
7. Strip steak—cooked to 75°C and pressed
6 1. To evaluate samples in the booths and begin independent evaluations; introduce evaluation in 2 sessions with a break. Session number 1: four steaks that vary in juiciness, MF tenderness, or CT amount based on previous training
Break: 15–20 min
Session number 2: four steaks that vary in juiciness, MF tenderness, or CT amount based on previous sessions
  • Abbreviations: CT, connective tissue; MF, muscle fiber.

  • Continue with training by increasing the number of steaks per session to 6. Each day, evaluate and review scores. Add exercises needed to increase understanding and repeatability of attributes. This example is for an experienced panel that is being refreshed. For a new panel, these exercises would be expanded and presented in more sessions depending on the panelists’ responses.

Table 19.

Examples of training sessions for a beef-flavor descriptive attribute panel using the beef lexicon (Adhikari et al., 2011)a

Session Session Goal Suggested Exercises Used in the Session
1 Introduce scaling Present universal scale for flavor intensity (Meilgaard et al., 2025)
  • 2.0 soda flavor in saltless saltine cracker

  • 5.0 apple flavor in Mott’s apple sauce

  • 7.0 orange flavor in Minute Maid orange juice

  • 10.0 grape flavor in Welch’s grape juice

  • 12.0 cinnamon flavor in Big Red chewing gum

Introduce basic tastes and recognize intensity levels across attributes using solutions from Meilgaard et al. (2025) for salt, sweet, bitter, and sour
Taste nonmeat items for flavor intensity and basic tastes from Meilgaard et al. (2025)
  • Lay’s Classic potato chips: rate potato flavor using universal scale; salt = 12.0; sweet = 4.5; sour = 1.5; and bitter = 2.0

  • Minute Maid orange juice frozen concentrate, reconstituted: rate orange flavor using universal scale; sweet = 8.0; sour = 3.5; and bitter = 1.5

  • Haagen-Dazs vanilla ice cream: rate vanilla flavor using universal scale; sweet = 12.0; salt = 2.0; sour = 2.0; and bitter = 1.0

2 Review universal scale and basic tastes Repeat exercises from session 1
Train on attribute 1: beef flavor, and aroma ID Provide definition and references for beef flavor and aroma ID
Panelists will evaluate individually and come to consensus
  • High beef flavor/aroma ID: prime top loin steak cooked on grill to 70°C

  • Low beef flavor/aroma ID: standard top loin steak cooked on grill to 70°C

3 Review beef flavor/aroma ID and have references available; continue to anchor using universal scale and basic taste references; introduce warm-up sample concept; introduce brown/roasted flavor/aroma attribute Have references for universal scale, basic tastes, and beef flavor/aroma ID available; panelists work through references at their own pace to anchor on attributes
Present warm-up sample: Choice top loin steak, aged 14 d, grilled to 70°C
Panelists evaluate beef flavor/aroma ID and basic tastes; discuss and come to consensus
Introduce definition and references for brown/roasted flavor and aroma
Present 3 samples and evaluate for all attributes introduced to this point
  • Choice top loin steak cooked to 125°F (57°C) for low brown/roasted flavor

  • Choice top loin steak cooked to 175°F (79.4°C) for high brown/roasted flavor

  • Choice top loin steak cooked to 70°C: unknown level; panelists will determine

4 Retrain on using previous exercises if panelists are having difficulty Present any previous references as needed to anchor panelist (Note: do this at the beginning of all sessions as needed; this will not be repeated for future sessions, but should be considered as an exercise for any session where panelists are showing inconsistencies in scaling for an attribute)
Present definition and references from lexicon for bloody/serum-like flavor/aroma attribute
Present definition and references from lexicon for metallic flavor/aroma attribute
Introduce bloody/serum-like and metallic flavor/aroma attributes Present 2 samples for panelist evaluation of all attributes introduced to this point
  • Select tenderloin steak grilled to 70°C: high bloody/serum-like and metallic attributes

  • Select top sirloin steak grilled to 70°C: unknown; panelists determine and come to consensus

5 Introduce fat-like flavor/aroma Present definition and references from lexicon for fat-like flavor/aroma attribute
Introduce liver-like flavor/aroma Present definition and references from lexicon for liver-like flavor/aroma attribute
Present 2 samples for panelist evaluation of all attributes introduced to this point
  • Prime strip steak broiled till 70°C: should have strong fat-like flavor and low liver

  • Cow strip steak broiled to 135°F (57°C): strong liver notes

6 Introduce green-hay-like flavor/aroma Present definition and references from lexicon for green-hay-like flavor/aroma attribute
Present definition and references from lexicon for umami flavor attribute
Introduce umami flavor Present 1 sample for panelist evaluation of all attributes introduced to this point
  • Prime grass-fed steak broiled to 135°F (57°C): green and umami notes

7 Introduce overall sweet flavor Present definition and references from lexicon for sweet flavor attribute
Introduce sweet aroma Present definition and references from lexicon for sweet and sour aroma attribute
Introduce sour aroma Present 2 samples for panelist evaluation of all attributes introduced to this point
  • Choice top loin steak grilled to 70°C

8 Calibration day Panelists will individually evaluate 3 muscles for the MAJOR NOTES
Warm-up: Select strip steak grilled to 70°C
Panelists will come to consensus
  • Sample 1—extra knuckle roast: no consensus, individual evaluation

  • Sample 2—extra top sirloin steak: no consensus, individual evaluation

  • Sample 3—extra eye of round roast: no consensus, individual evaluation

Determine panel proficiency on MAJOR NOTES
9 Overview major notes Present definition and references from lexicon for animal hair aroma attribute
Present definition and references from lexicon for barnyard aroma/flavor attribute
Introduce animal hair aroma Present 3 samples for panelist evaluation, all attributes introduced to this point
  • Tenderloin from a bull broiled to 74°C for animal hair aroma and barnyard flavor/aroma

Introduce barnyard aroma/flavor
  • Select top loin steak grilled to 70°C

  • Choice knuckle roast, roasted to 70°C

10 Review major notes Present definition and references from lexicon for animal hair aroma attribute
Introduce burnt aroma/flavor Present definition and references from lexicon for barnyard aroma/flavor attribute
Introduce rancid aroma/flavor Present 4 samples for panelist evaluation of all attributes introduced to this point
  • Standard strip grilled to greater than 175°F (∼79.4°C) or until visibly burnt

  • Low Choice strip stewed to 155°F (∼68°C)

  • Select inside round roast (70°C)

  • Choice flat iron steak (70°C)

11 Review major notes Present definition and references from lexicon for heated oil aroma attribute
Present definition and references from lexicon for chemical aroma/flavor attribute
Introduce heated oil aroma/flavor Present 3 samples for panelist evaluation, all attributes introduced to this point
  • Top Choice top butt broiled to 155°F (∼68°C): chemical aromas and flavors

Introduce chemical aroma/flavor
  • Choice eye of round roast (70°C)

  • Choice tenderloin (70°C)

12 Review major notes Present definition and references from lexicon for leather aroma attribute
Present definition and references from lexicon for apricot flavor attribute
Introduce leather (old) aroma Present 3 samples for panelist evaluation of all attributes introduced to this point
  • Cow top round stewed to 165°F (∼74°C): leather aromas

Introduce apricot flavor
  • Select tenderloin (70°C)

  • Choice bottom round roast (70°C)

13 Review major notes Present definition and references from lexicon for green aroma/flavor attribute
Introduce green aroma/flavor Present definition and references from lexicon for asparagus flavor attribute
Introduce asparagus flavor Present 4 samples for panelist evaluation of all attributes introduced to this point
  • Cow strip broiled to 135°F (∼57°C): green flavors

  • Cow strip stewed to 145°F (∼63°C)

  • Select knuckle roast (70°C)

  • Select flat iron steak (70°C)

14 Review major notes Present definition and references from lexicon for green aroma/flavor attribute
Present definition and references from lexicon for asparagus flavor attribute
Introduce musty-earthy/humus aroma Present 4 samples for panelist evaluation, all attributes introduced to this point
  • Cow tenderloin roasted to 145°F (∼63°C)

Introduce cumin aroma/flavor
  • Choice inside round roast (70°C)

  • Select top sirloin (70°C)

15 Review major notes Present definition and references from lexicon for floral aroma/flavor attribute
Present definition and references from lexicon for beet aroma/flavor attribute
Introduce floral aroma/flavor Present 4 samples for panelist evaluation of all attributes introduced to this point
  • Cow tenderloin stewed to 175°F (∼79.4°C): beet flavors

Introduce beet aroma/flavor
  • Select bottom round roast (70°C)

  • Choice strip steak (70°C)

16 Review major notes Present definition and references from lexicon for floral aroma/flavor attribute
Present definition and references from lexicon for beet aroma/flavor attribute
Introduce medicinal aroma Present 4 samples for panelist evaluation of all attributes introduced to this point
  • Bull tenderloin grilled to 145°F (∼63°C): chocolate aroma

Introduce chocolate/cocoa aroma/flavor
  • Choice top sirloin steak (70°C)

  • Select eye of round roast (70°C)

17 Review major notes Present definition and references from lexicon for charcoal aroma/flavor attribute
Present definition and references from lexicon for wood aroma/flavor attribute
Introduce charcoal aroma/flavor Present definition and references from lexicon for spoiled-putrid aroma attribute
Present 4 samples for panelist evaluation of all attributes introduced to this point
Introduce wood aroma
  • Smell spoiled standard tenderloin grilled to 155°F (∼68°C): spoiled-putrid aroma

  • Prime strip steak food service gas grilled to 155°F(∼68°C): smoky charcoal aroma

Introduce spoiled-putrid aroma
  • Select strip (70°C)

  • Choice strip (70°C)

18 Review major notes Present definition and references from lexicon for dairy aroma/flavor attribute
Introduce dairy aroma/flavor Present definition and references from lexicon for buttery aroma/flavor attribute
Introduce buttery aroma/flavor Present 4 samples for panelist evaluation of all attributes introduced to this point
  • Top Choice top butt roasted to 145°F(∼63°C): dairy aroma

  • Bull tenderloin grilled to 145°F(∼63°C): buttery flavor

  • Select eye of round roast (70°C)

  • Choice eye of round roast (70°C)

19 Review major notes Present definition and references from lexicon for milk aroma/flavor attribute
Present definition and references from lexicon for sour milk/sour dairy aroma/flavor
Present 3 samples for panelist evaluation of all attributes introduced to this point
Introduce milk aroma/flavor and sour milk/sour dairy aroma/flavor
  • Milk-fed veal strip steak cooked milk

  • Select top sirloin steak (70°C)

  • Choice top sirloin steak (70°C)

20 Review major notes Present definition and references from lexicon for stale aroma/flavor attribute
Present definition and references from lexicon for soapy aroma attribute
Introduce stale aroma/flavor Present definition and references from lexicon for warmed-over aroma/flavor attribute
Present 4 samples for panelist evaluation of all attributes introduced to this point
Introduce soapy aroma
  • Cow strip broiled to 135°F (∼57°C): refrigerator stale notes and warmed-over flavor

  • Select tenderloin broiled to 175°F (∼78°C): warmed-over aroma and flavor

Introduce warmed-over aroma/flavor
  • Choice flat iron steak (70°C)

  • Select flat iron steak (70°C)

21 Calibration: evaluate all attributes Warm-up: Choice strip steak
  • Sample 1: Select eye of round roast

  • Sample 2: Choice top sirloin steak

  • Sample 3: Choice knuckle roast

  • Sample 4: Select flat iron

22 Calibration: evaluate all attributes Warm-Up: Choice eye of round roast
  • Sample 1: Select inside round roast

  • Sample 2: Choice bottom round roast

  • Sample 3: Choice flat iron

  • Sample 4: Select tenderloin

  • Abbreviation: ID, identification.

  • Training schedule is adjustable and may be changed to optimize panel performance, as the level of previous training may influence the ability of panelists to complete each exercise.

With ground beef patties, the following processes provide excellent variations in sensory properties for training. Potential sensory training sessions for ground beef would follow similar logic to the information presented in Tables 8 and 18, but with the addition of respective texture attributes. Some suggested variations to include are: variation in grind size; fast vs. slow freezing; use of added ingredients that alter flavor and texture, such as gels, gums, Textured Vegetable Protein (TVP); different fat levels; and different cookery methods (charbroiling versus microwave).

During the early stages of training, the panel leader should try to identify the extremes and the middle of the rating scale. It is necessary to refer to some standards, such as the psoas major muscle, for extremely tender and hot-boned longissimus or semitendinosus from an old cow carcass chilled in ice water for extremely tough samples. As training progresses, the panelists should be able to identify other points along the rating scale. Once panelists have begun to scale for individual attributes, training sessions that include the evaluation of a combination of 2 attributes should be conducted. Successive additions of other attributes should occur as panelists illustrate knowledge and confidence in their abilities as the complexity of the evaluation increases (Figure 24).

The panel leader will need to provide very specific instructions to panelists regarding procedures to follow in measuring the various sensory attributes during chewing. With the complexity of added ingredients now being used in products, especially processed meat products, additional training sessions would be needed to ensure consistent understanding and scaling of each attribute. For example, with ground beef, overall tenderness ratings might not provide sufficient information on texture differences. Texture attributes, as presented in Table 8, most likely describe texture differences across patties. During training, the use of gels and gums might create softness in the product, but also a tacky or cohesive property. These experiences may not reflect the ground beef that will be presented during testing, but will provide a wider array of experiences for panelists. For texture attributes, stating which teeth should be used as well as the number of chews should be standardized when possible. Positioning of the sample on the teeth, such as orientation of the muscle fibers for whole-muscle products, or biting down with molars on a standardized portion of a chicken nugget or sausage patty from top to bottom, should be addressed. Molars can be used to detect firmness, compression, and rate of breakdown, while incisors provide information relative to shear. Incisors help detect connective tissue and detect the rubberiness or elasticity of a small piece of connective tissue. The tongue can also help detect connective tissue by pushing it through the chewed mass. Panelists must be instructed to perform a considerable sample breakdown, often employing a high number of chews (ready to swallow) before assigning a score for the volume of residual connective tissue present. The presence of surface crust or breading on cooked samples might present unique problems relative to sample breakdown during chewing and uniformity of chewed pieces. In studies where designed differences exist in fat content, juiciness should be evaluated early (initial) and later (sustained) in the chewing process.

Training time is a function of product testing method and procedural variables, and it increases with the number of attributes to be studied. The availability of panelists and the responses of panelists to training influence training time. Individuals should be evaluated during training and while the study is in progress. Delays or interruptions of more than 2 wk in a series of tests should be followed by refresher training sessions and perhaps an evaluation. Training is never completed, and day-to-day variation among panelists is an issue that must continually be monitored.

Trained descriptive attribute panels should not be asked to evaluate any attribute in terms of like/dislike or acceptability. Those responses should be obtained only from a consumer panel (see Consumer Panels section). After the panel has been trained, ballot development sessions begin with the standard lexicon and are enhanced with attributes unique to a given study. References for attributes can be found in the scientific literature and Meilgaard et al. (2025). Munoz and Civille (1998) discussed attribute scaling and the development of common lexicons for use in descriptive analysis in ballot development sessions.

Performance evaluation for descriptive attribute panels

Performance evaluation can begin soon after training starts, throughout training, and at the end of training to measure each panelist’s ability to (1) discriminate among samples, (2) rate samples consistently each time (repeatability), and (3) rate samples similarly to other panel members (agreement). These key measures should be monitored as training progresses to identify sources of any poor performance, such as inability to use the scale correctly to rate attribute intensity (scale use) and failure to use attributes similarly to other panelists (lexicon use) (ASTM Standard E3000-24, 2024).

To conduct a performance evaluation, 9 samples are selected to cover the full range of test attributes. Panel evaluation is spread over 4 d with 3 sessions per day and 3 samples per session. The design is outlined in Table 20. Samples can also be presented in a Williams’ square arrangement (Williams, 1949) to permit each treatment to appear in every session as well as every serving position within the sessions (Table 21) or as a balanced lattice design (Table 22). Data analysis can be conducted using individual panelist data for each attribute evaluated as a 1-way ANOVA with 9 treatments and 4 observations per cell. From the ANOVA table, the F-ratio (F = Mean Square treatments/Mean Square error) is calculated. The F-ratio is a measure of a panelist’s ability to discriminate while being able to repeat ratings on the same sample on different days. The degree to which a person discriminates among samples and is consistent in replicated judgments is reflected in the F-ratio (Cross et al., 1978). A larger F-ratio indicates a better-performing panelist. Candidates can be ranked on the basis of their F-ratios, which are product, attribute, and study dependent. The sensory panel leader can identify attributes and panelists with low F-ratios for further training.

Table 20.

Example of how to randomize 9 samples to 3 sessions over 4 sensory days for performance evaluation using a random numbers table

Day1 Day 2 Day 3 Day 4
Session Session Session Session
1 2 3 1 2 3 1 2 3 1 2 3
S9 S5 S6 S8 S6 S3 S2 S6 S7 S6 S8 S4
S8 S3 S1 S1 S9 S7 S4 S3 S1 S7 S3 S2
S2 S4 S7 S5 S2 S4 S9 S5 S8 S1 S5 S9

    Abbreviation: S, sample number.

Table 21.

Example of a 6 × 6 Williams square adapted from Williams (1949)

Serving Order
Sensory Session 1 2 3 4 5 6
1 A B F C E D
2 B C A D F E
3 C D B E A F
4 D E C F B A
5 E F D A C B
6 F A E B D C
Table 22.

A 3 × 3 balanced lattice design using 9 samples or sensory treatments assigned to 12 sessions, where 3 sessions are conducted per sensory day over 4 sensory days. Adapted from Srisuradetchai (2012)

Sensory Day Session Samples Served Sequentially/3 per Session
1 or replicate 1 1 1 2 3
2 4 5 6
3 7 8 9
2 or replicate 2 1 1 4 7
2 2 5 8
3 3 6 9
3 or replicate 3 1 1 6 8
2 2 4 9
3 3 5 7
4 or replicate 4 1 1 5 9
2 2 6 7
3 3 4 8

The data also could be evaluated using a 2-way ANOVA with one observation per cell. The design can be treated as a balanced lattice design (Cochran and Cox, 1957; Srisuradetchai, 2012). With this design, day and session effects can be studied. Srisuradetchai (2012) provided R and Statistical Analysis Software (SAS) code for this analysis.

The data can be analyzed where the effect of panelist, treatment, and panelist by treatment interaction is tested. In this analysis, individual data for each panelist for the attributes evaluated are used. The experimental unit is defined as a panelist’s evaluation of one sample within a session and sensory day for an attribute. In the example in Table 20, if 10 panelists were being evaluated, there would be 10 panelists who evaluated 9 samples across 4 sensory d, or a total number of 360 observations in the dataset. The effects of sensory day, treatment, panelist, and panelist by treatment interaction can be determined. Order effects can be evaluated or used as random effects in the model. By using ANOVA, the residual error will include unexplained differences between samples. Least-squares means can be calculated for each main effect and the interaction in the model. This enables the panel leader to understand variances associated with these effects, where panelists are in relationship to other panelists, if panelists are evaluating samples differently across days, and if panelists rank treatments differently. From these results, exercises can be set up to address inconsistencies in panelist performance.

The number of descriptive panelists selected should be based on performance evaluation results. Performance evaluation should be conducted intermittently during training and at the end of training. A person whose evaluation is less than satisfactory should not be used as a panelist just to achieve a predetermined panel size. According to ASTM MNL90-2ND-EB (2023), the ideal panel size is 12–15, with 10–12 being used in any given study, which is a good general rule. Having fewer than 10 on any given study can result in too much dependence upon any one individual’s response. However, fewer panelists (5 or more) can be used if they are sufficiently trained, have extensive experience (greater than 200 h of training), and have been validated. Labbe et al. (2004) evaluated panelists’ performance before training and twice during training. They found that training affected absolute scores and variation across panelists. Additionally, training improved the discriminating ability of panelists. The number of attributes or the complexity of the products being evaluated also impacts the level and length of training. Chambers et al. (2004) evaluated the effect of short, moderate, and extensive levels of training (4 h, 60 h, and 120 h, respectively). They found that panelists with 4 h of training detected differences in many texture and some flavor attributes. However, with extensive training, variation between panelists was reduced, differences were detected in a greater number of attributes, and there was increased ability to detect differences. Therefore, the number of panelists used in a study is dependent on the level and length of training, the number of attributes, the complexity of the samples, the validation of the panelists, and the consistent monitoring of panelist performance.

Before conducting a study, the expected sensitivity or level of discrimination of the panelists should be identified. For example, panelists may be able to determine differences at of 2, 1, or 0.5 using a 16-point scale. During panel validation, the panel leader can use the predetermined level of sensitivity and examine the root mean square error and differences in least-squares means to determine if the panelists are able to discriminate at this level. If panelists have not reached the desired level of sensitivity, additional training is required. Chambers et al. (2004) showed how the length of training improved sensitivity and reduced variation across panelists. By setting the desired level of discrimination prior to initiation of training, a clear understanding of when training is sufficient can be obtained so that testing can proceed.

The panel leader needs to carefully evaluate repeatability and panel agreement based on the variability in responses. As highly trained and experienced panelists will be better at discrimination and agreement, and will be more repeatable, smaller panels with a minimum of 5 to 8 panelists can be used. Keep in mind, though, that it is more important to have panelists who can discriminate and have fewer numbers than to have higher numbers of panelists who are not adequately trained. Chambers et al. (1981), Chambers et al. (2004), and Djekic et al. (2021) provide insight into the benefits of training and the number of panelists needed to collect sensory data. The decision of who should or should not be on the panel should be based solely on evaluations during training, at the end of training, during a study (especially if the study has a long duration, such as ≥4 wk), and at the end of the study during data analysis. Labbe et al. (2004) showed that subsequent evaluations of panelists’ performance are necessary throughout a study.

Training for new panels could last 2 to 6 mo, depending on the level of experience of individuals involved, the number of attributes being evaluated, and the descriptive analysis method being used. For example, panelists with over 200 h of training and multiple years of experience within a product or across different products may only need 10 h or less of training for a new study that uses experimental units with which they have experience. However, naive or new panelists may require 3 to 6 mo of training to approach the same level of sensitivity. Refer to Cross et al. (1978) for review and validation of this technique, or see the discussion of panelist effects under the data analysis section of these guidelines.

Monitoring panelist performance and maintenance

Performance records should be maintained for each panelist and should be periodically reviewed by the panel leader. Overall panel performance evaluations (ANOVA that includes panelist effects and panelist by treatment interactions) alert the panel leader to any problems that individual panelists may have. Consistently poor performance may indicate that the panelist was incorrectly selected for the study or that physical or psychological distractions are preventing them from making sound evaluations. For panelists with poor performance identified during or between studies, additional training sessions should be conducted. Their data should only be used once they have passed the subsequent performance evaluations.

The importance of cooperation and performance as it relates to the entire project should be stressed. Feedback to the panelists as to their importance to the total research program is an integral part of keeping the group willing to participate. Some institutions have successfully used paid panelists. A system of rewards can also be instituted. For example, cookies, candy, cake, or ice cream can be used. Those who have used paid panelists have found them to be more highly motivated and more consistently available than “in-house” personnel. As a result of being on the sensory panel, panelists may develop other outside activities together. Holding periodic social events among management and panelists may help create the stimulus to stay with the group. Care should be taken to ensure outside activities and discussions are not brought into the sensory evaluation environment. Attributes of test products mustn’t be discussed during outside social activities until after an experiment is completed.

Panelists should be continually trained to maintain relevant attributes fresh in their memory. Reference samples should be available. When necessary, screen other potential panel members to maintain or enlarge the panel. In any particular study or test, substitution of panel members should be avoided. Refresher training should be conducted before starting a new experiment and after an extended interruption of testing.

There is always a concern that panelists may vary over evaluation days. Warm-up samples, where panelists evaluate a sample independently and then discuss the sample and individual attributes to come to a consensus, are a practice to help reduce day-to-day variation. In large studies, after sufficient data are collected, panelist, day, treatment, panelist by treatment, and panelist by day effects can be evaluated. Note that the effect of treatment and panelist by treatment effects will have limited information, but including them in the analysis removes these effects from the error term and allows for more concise evaluations of panelist and day effects. This information can then be used to direct subsequent training sessions and provide direction for altering warm-up samples to reduce the Error of Habituation. This is where the same response is given for an attribute that slowly changes in intensity. If the Error of Habituation is apparent for warm-up samples, warm-up samples can be changed to induce variation, or a second warm-up may be used to maintain panelist scaling and to reinforce the use of scales.

Another method of reducing panelist effects and variation is to supply references to panelists to review before conducting a warm-up sample, and then to have key references with each panelist during product evaluation. During training, the use of references to anchor panelists should be emphasized. When discrepancies are found, panelists and the panel leader should use references to anchor the panelists for that particular attribute.

Consumer sensory panels

While the use of a trained panel to assess treatment differences in a set of products is critical in determining if the treatment differences are detectable, it is also critical to know if these differences impact consumer acceptance. Are these differences large enough for consumers to detect? Are they important enough to affect overall acceptance? Consumer testing is often a key part of a study, as the fate of a food product rests on acceptance by the consuming public. It is only through careful planning of the consumer test, including choosing the most appropriate test method, that the researcher can get reliable, repeatable results that can be projected to the real consumer population.

Trained and consumer panels can both be useful to determine if differences exist within and between products. Consumer tests should be used if it is important to determine if consumers can detect the differences in the product due to the addition of some treatment. In this situation, Triangle or Tetrad tests can be conducted with consumers to determine if the consumers detect a difference between the products. An example of this type of test would be the addition of a microbiological intervention to ground beef and its impact on the sensory properties of the product. As consumers are less likely to detect differences compared with trained panelists, it is advisable to use consumer Triangle or Tetrad test data in conjunction with trained panel data when making decisions on whether or not there is a perceivable difference between the products.

Additional points to consider when designing and conducting a consumer sensory test are experimental design, protocol design, test method, ballot development, moderator selection, and interaction with the sensory subjects. Meilgaard et al. (2016, 2025), Lawless and Heymann (2010), ASTM MNL26-3RD-EB (2020), and Stone and Sidel (2004) are excellent references for information on the design of consumer panels.

Determining the test type

The first step in setting up a consumer study is determining the most appropriate test method to use. The test method should be selected based on the study objectives and on how the results will be used. Consumer panels can be defined as either qualitative or quantitative tests.

Qualitative tests provide a subjective response of a sample from consumers by having those consumers talk about their feelings regarding the sensory properties of a set of products (Meilgaard et al., 2025). Examples of qualitative tests include focus groups or panels, mini groups, dyads and triads; one-on-one interviews; and workshops. Quantitative tests are used to determine consumers’ sensory perception of products by using a set of questions to measure preference, liking, and impressions of various sensory attributes. Quantitative tests can be conducted using CLT with pre-recruited participants; nonprerecruited participants, such as a mall intercept test; or using HUT.

The CLT test is where consumers are brought into a controlled environment to evaluate products. In these situations, sample preparation, environment such as temperature, noise, distraction, lighting, and sample size can be controlled, and there is some assurance that these factors have minimal effect on sensory verdicts. However, the CLT environment is artificial, does not include the influence of the home or restaurant environment, and excludes the influence of family or other household members. The HUT tests allow consumers to interact with the packaged product, cook the product, season as desired, and consume the product normally. The influence of others may influence the perception of the product. Variation is greater across and within treatments in HUT versus CLT consumer studies. However, information on how consumers prepare, season, what recipes are used, cooking methods, degree of doneness, and other factors can be assessed in HUT.

When using mall, store, or restaurant intercept studies, the number of questions and time needed to conduct the study should be limited. When consumers are intercepted, they may have a limited attention span due to time restraints for the task that they were completing when intercepted.

If the intent of the consumer evaluation is to determine which product is preferred, preference testing should be performed. How well a product is liked by consumers can be ascertained through acceptance tests. Hedonic scales (like/dislike) are used to indicate the degree of acceptability and whether preferences exist among products. JAR scales and intensity scales can be used to determine when an attribute is too high/strong or too low/weak.

Protocol design

The number and type of panelists, as well as sample preparation and presentation, can significantly influence consumer responses and, therefore, significantly affect or limit the interpretation of the results from the study.

There are several factors to consider when determining the optimal panel size, including test objectives such as product reformulations for improvement, similarity in an ingredient substitution test, or testing against competition for claims testing. The objective should define if the test is one-sided (Is the test product liked more than the current) or 2-sided (both samples are equally viable choices). The test method should be defined as a CLT, HUT, or other method. Acceptable levels of risk must be predetermined by defining the appropriate α level (probability of committing a Type I error) and β level (probability of committing a Type II error). The size of the expected or meaningful difference should be estimated. For example, differences of 0.4 or 0.8 units on a 9-point hedonic scale may be important. The number of locations that will be used for testing should be determined, as well as whether or not any subgroups will be evaluated to ensure that representative populations are used in the study.

According to ASTM MNL26-3RD-EB (2020), 100 consumers are usually adequate for most consumer tests, but the optimum number is influenced by the factors listed above. In addition to α and β levels and the magnitude of difference to be detected, Hough et al. (2006) discussed determining the optimum panel size based on the standard error of the experiment. The standard error of the experiment is not known beforehand but can be estimated from previous studies. For consumer studies, data from past studies can be analyzed by analysis of variance, and the standard error of the experiment can be determined by the square root of the mean square error (Hough et al., 2006). Alpha and β are key measures associated with the risks of committing Type I and Type II errors, respectively. The levels are determined based on the test objectives and influence the sample size. If the objective of the study is to determine whether or not the test product is significantly better in flavor than the control, then Type I error is the critical measure. In this case, a Type I error is the risk in declaring that the test product has a better flavor, although, in reality, the test product does not. Conversely, if the objective is to determine whether or not the test product is liked as well as the control after an extended storage period, Type II error is the critical measure. The risk is declaring that the test product is equally acceptable to the control. In reality, the test product is inferior. Larger sample sizes are typically required for parity tests (Type II error focus) than for superiority tests (Type I error focus).

Researchers also must consider the testing conditions. Under controlled conditions in sensory panel booths, there are fewer distractions, so fewer panelists are needed compared with a less-controlled setting, such as an HUT. The sample size will also need to be increased if testing in multiple locations because regional differences in acceptability might exist. A minimum of 50 consumers per location is generally recommended to test for regional differences. In situations where consumer segments exist, such as preferences for spice level, the sample size will also need to be increased to allow for data segmentation.

Review and evaluation of previous studies with a statistician is beneficial for establishing the appropriate number of consumers because consumer testing can be expensive. Additional references for sample size determination can be found in ASTM Standard E1958-22 (2022) and Lawless and Heymann (2010).

Panelist selection

The application of results will depend extensively on the criteria used in selecting test participants. Selecting the proper panelists is key to obtaining test results that can be projected to the target population. When selecting panelists for a consumer test, the following factors should be considered. Consumers in the target population should be selected as panelists to get an accurate measure of consumer acceptance. The use of employees should be limited to product guidance in the early product development phase or for low-risk product tests, i.e., lower volume products. When using employees as consumer panelists, limit the number of times they can participate so that they don’t become overly sensitive to product differences. Additionally, do not use employees who are knowledgeable about the study. It is also best to avoid Research and Development and Quality Assurance employees because of their technical backgrounds and relationship with the product. Consumers who use the product can provide data that are more reflective of real-life evaluations and an understanding or verification of acceptance, especially with high-risk or high-volume products.

Product use patterns (i.e., heavy versus light users) may be used when selecting consumer panelists. Preferences for the degree of doneness are also critical because serving a medium-rare meat sample to a panelist who typically eats their meat well done would not give a true read on consumer acceptance. While the demographics of the target population, such as age, gender, income, education, and household size, are often used to help identify the target consumer, they do not always guarantee that they are actual product users. If product use is critical, the screener must include questions on the frequency of use. The number of markets or locations that need to be tested needs to be considered. Which cities have the largest number of the target consumers, and are there regional preferences in the product preferences that should be taken into consideration? These factors influence the consumer sensory outcome and interpretation of results.

Ballot development

The consumer sensory ballot, which is defined as the sensory instrument for consumer testing, is critical to ensure that the sensory professional is accurately testing consumer responses and not biasing the responses with a poorly designed ballot. In designing a ballot, some general rules are to ask for responses in the order in which they are normally encountered while eating or sampling the product. For example, most consumers will evaluate the appearance, then flavor, and lastly texture. A combination of close-ended (i.e., hedonic scales) and open-ended (no defined response) questions should be used. Scales should be consistent with the ballot for the number of points of determination, scaling, and style. For example, if a 9-point hedonic scale is being used, do not change to a 5- or 7-point hedonic scale across questions within the same ballot. It is appropriate, however, to include intensity (5-, 7-, or 9-point) and/or JAR (typically 5-point) scales for impressions of specific attributes. When using JAR scales, use penalty analysis to determine the impact of an attribute not being JAR. Do not reverse the end-anchors on the scales within one ballot. Provide clear, concise directions and unambiguous questions to ensure that the test accurately measures consumers’ sensory perceptions and keeps questions to a minimum. When selecting the number of points on a scale, a decision has to be made to have either an even or an odd number of points. Odd-numbered scales provide a neutral point that can be anchored or not anchored. Even point scales do not have a neutral point. In some studies, the researcher may need to understand if the consumer is neutral or does not have a defined like or dislike of the product, or it may be important to force the consumer to make a decision for like or dislike. Both types of scales are used, but the researchers should understand how the scale impacts consumers’ thinking and perception.

The most important consumer question is almost always the consumer’s overall like/dislike of the product. The position of this question can influence the answer. When asked about the appearance and taking the first bite of the product, the answer reflects the consumer’s first impression. When asked at the end of the ballot after questions concerning the appearance, juiciness, tenderness, and flavor, the answer reflects the consumer’s first impression and any influence of the questions asked. The researchers must consider this when designing and interpreting the results.

Hedonic scales

The hedonic scale can be shown horizontally or vertically, and scales can be verbally anchored at each point along the scale with 9 categories or as continuous line scales (Figure 28). When using the hedonic scale horizontally, it should be anchored with “Dislike Extremely” on the left and “Like Extremely” on the far right (Figure 28). The purpose of anchoring the scale at each point is to encourage a continuum with equal spaces between each successive increase in like/dislike or preference.

Figure 28.
Figure 28.

Examples of commonly used hedonic and just-right scales.

Meilgaard et al. (2025) and Lawless and Heymann (2010) give examples of hedonic scales that can be used in consumer testing. These scales range from the anchored, 9-point hedonic, to the end-anchored, 9-point hedonic; to the end- and neutral-anchored, 9-point hedonic; to the non-balanced scale (more or less categories of like in relation to dislike). Many different forms of hedonic scales can be used without major effects on the value of the results, as long as the essential feature of verbal anchoring of clearly successive categories is retained. There should be at least 5 categories. Replacement of the verbal categories with caricatures representing degree of pleasure and displeasure (smiley scale) has been used, but studies on this type of scale have indicated potential issues with how children interpret the caricatures, so these scale types are not recommended (ASTM Standard E2299-13, 2021). If testing with children, a 9-point “super good/super bad” scale developed by Kroll (1990) has been used successfully. It can be truncated to 7 points for younger children if needed.

Just-about-right and intensity scales

Intensity scales and JAR scales can be used to determine if products differ significantly in the levels of specific attributes. These scales can be bipolar or unipolar. Intensity scales can be line scales or category scales, and they can be end-anchored or fully anchored. When used as category scales, intensity scales are often 5-point, 7-point, or 9-point. If using a mixture of JAR and intensity scales on the ballot, it is good practice to use different scale lengths for the 2 scale types to minimize confusion about what is being asked.

JAR scales are typically category scales and are most often 5-point, with 7-point and 9-point scales used less often (Figure 28). Because an attribute can vary without negatively impacting consumer liking, it is critical when using these scales that penalty analysis be used to assess the impact or penalty of an attribute not being JAR. Penalty analysis links the drop in overall liking to the proportion of respondents rating an attribute either too high or too low. It, therefore, provides the ability to prioritize product optimization opportunities by determining the critical product attributes that penalize product acceptance the most. It is important to note that penalty analysis should only be run for JAR attributes with 20% above or below JAR. Running the analysis with less than the 20% cutoff results in too much dependency on any given respondent. For studies with smaller sample sizes or when analyzing subgroups within a study, the cutoff should be increased to at least 25%. Refer to ASTM MNL63 (2009) for additional information on penalty analysis.

Several SAS packages, such as XLSTAT, offer an automated penalty analysis feature. If unavailable, however, the penalty for an attribute not being in the JAR is calculated in the following manner:

  • Collapse JAR scales to 3 points.

    • o

      Too little (TL)

    • o

      JAR

    • o

      Too much (TM)

  • For each attribute with 20% or more of respondents rating it either too high or too low, calculate the mean drop in overall liking for not being JAR.

    • o

      Mean dropTL = AverageJAR – AverageTL

    • o

      Mean dropTM = AverageJAR – AverageTM

  • Calculate the total penalty for each attribute not JAR.

    • o

      Total penalty = (Mean drop) * (% not JAR)

Penalties can be shown graphically in 2 ways, as shown below in Figures 29 and 30. Total penalties can be categorized into different tiers based on potential opportunity to improve overall liking if the offending attribute is brought to JAR: < 0.25 = low-tier penalty: no changes needed; 0.25 to < 0.50 = mid-tier penalty: consider changing; and ≥ 0.50 = high-tier penalty: must change.

Figure 29.
Figure 29.

Mean drop in overall liking as a function of the percent not at just-about-right ratings for the various attributes. Abbreviation: JAR, just about right.

Figure 30.
Figure 30.

Total penalties of the various attributes.

In the above example, results showed that 35% of the panelists rated the sample “too salty,” but the drop in mean overall liking was 0.5 units, which resulted in a total penalty of 0.18 units (Figure 29). This penalty is considered low tier, indicating the salt level does not need to be adjusted. Additional information on penalty analysis and the uses and abuses of JAR scales can be found in ASTM MNL63 (2009).

An example of a consumer ballot is presented in Figure 31 using a combination of hedonic, intensity, and JAR scales. Note that demographic and product use questions can be included on a consumer ballot. Computerized demographic and use questions are used on prescreening questionnaires so that these questions do not influence consumers’ expected or actual knowledge of the project objectives (defined as expectation error). Instructions are also included on the ballot, even though a moderator should be present during testing to assist participants in understanding what is expected of them.

Figure 31.
Figure 31.

An example of a consumer sensory ballot using 9-point hedonic scales and 5-point just-about-right scales.

Instrumental Measures of Textural Properties and Tenderness

Instrumental methods are used to quantify/measure mechanical properties that can be related to valued sensory properties. Instrumental methods can provide an assessment of product tenderness or texture. In general, measures of tenderness are more related to intact muscle cuts (e.g., steaks, chops, breasts) than further-processed meats. Texture represents a broader category of physical properties that can be measured on a variety of ground, processed, and comminuted meats. Instrumental tenderness/texture methods fall into 3 categories that include: empirical, imitative, and fundamental methods. Empirical methods and imitative methods do not provide independent mechanical properties based on material science principles. Imitative methods are designed to mimic some aspect of how an individual chews the product (mastication). A fundamental textural method is able to provide distinct, pure engineering mechanical properties such as stress and strain. Appropriate decisions need to be made to match the correct method with the type of meat or processed meat to be analyzed and the objective of the measurement.

It should be acknowledged that a trained sensory panel and instrumental measures only give measures of relative differences in tenderness and texture. They provide little indication of the acceptability of a given tenderness or texture measure, other than that derived from associating objective measures with consumer acceptance data. Acceptability of a given level of tenderness or texture can only be determined by the ultimate users, consumers (Garmyn, 2020). Historically, the lack of consumer data on meat tenderness and texture has led to over-interpretation of objective measures as indicators of whether meat would be considered tough, tender, or acceptable. Although much progress has been made in collecting meaningful consumer acceptance data, more data on the relationship between consumer impressions of meat tenderness/texture and objective measures of tenderness/texture are needed.

Methods used to assess instrumental tenderness in whole-muscle and processed meats are discussed. Table 23 summarizes recommended texture assessment methods for whole-muscle meat and processed meat products. Processed meats represent a broad category. For the purpose of selecting an appropriate tenderness or texture method, this category includes fresh ground meat and poultry products, such as patties, as well as more traditional processed products, including marinated meats, whole-muscle dried meats, sliced luncheon meats, and sausages. Table 23 presents suggested texture methods for processed meat products across classifications. It should be noted that some texture tests for processed meats may be conducted on raw and or cooked products. Researchers should review how texture measurements are applied in the scientific literature, particularly for raw processed meat products.

Table 23.

Suggested texture measurements for processed meat products

Product Classification Product Description Products Texture Methods
Whole intact muscle Meat and poultry products that maintain intact individual or multiple muscle complexes
  • Beef, pork, lamb, and veal whole meat cuts

WBSF, SSF, AKS, MORS, BMORS, star probe
  • Chicken breast and thigh meat

  • Fish filets

Whole intact muscle processed meats Meat and poultry products that maintain intact individual or multiple muscle complexes
  • Marinated chicken breasts

WBS, AKS, BMORS, star probe, torsion testing, penetrating probe/steel ball
  • Sectioned and formed processed deli meats (cured and smoked boneless ham, oven-roasted turkey breast, corned beef, pastrami)

Patties Seasoned coarse-ground meat and poultry products
  • Ground beef patties

AKS, TPA
  • Turkey burgers

  • Sausage patties

Sausage: finely comminuted Finely chopped/emulsified processed meats in which the cooked meat matrix is homogenous and isotropic
  • Frankfurters and bologna (links, unsliced)

TPA, torsion testing, puncture (e.g., star probe, blunt rod: skin assessment)
Sausage: coarse ground Sausages that are coarsely ground (∼3/8 to 1/8” plate)
  • Bratwurst

TPA, AKS
  • Polish sausage

  • Breakfast links

Sausages: sliced Sliced luncheon meat
  • Roast beef

Tensile testing for intact muscle processed meats
  • Ham

  • Bologna

Sausages: specialties Variety meats such as liver
  • Braunschweiger

Compression (single-cycle TPA)
Fish: filet/steaks Filets
  • Marinated catfish filets

Compression (single-cycle TPA)
  • Salmon steaks

Whole-muscle dried, processed meats Products still have intact muscle sections sufficient to access and are dried
  • Beef jerky

Tensile testing
Bacon Thin-sliced processed meats that maintain intact muscle sections
  • Bacon (pork belly)

Tensile testing, AKS
Pizza topping Small bite-sized processed meats used as toppings and/or inclusion into sandwich-type foods (tacos, wraps)
  • Sliced pepperoni

AKS
  • Ground seasoned sausage

  • Cooked hamburger

  • Abbreviations: AKS, Allo-Kramer shear force; BMORS, Blunt Meullenet-Owens razor shear; MORS, Meullenet-Owens razor shear; SSF, slice shear force; TPA, texture-profile analysis; WBSF, Warner-Brazler shear force.

Warner-Bratzler shear force (WBSF)

This measure can be obtained either with the original WBSF machine or with WBSF blade attachments to an automated testing machine (e.g., Instron, United, Texture Technologies, etc.). In addition to peak load (maximum shear force), other traits that might be useful can also be obtained with an automated testing machine. V-notch blades used for shear force should either be the blades made for WBS machines by G-R Manufacturing (Manhattan, KS, USA) or blades sold by the testing machine manufacturer (Figure 32). These blades are milled to exact specifications, including the bevel on the cutting edge. Unless in-house manufactured blades meet these exact specifications, they should not be used.

Figure 32.
Figure 32.

Examples of acceptable cores and the Warner-Bratzler shear force blade when conducting Warner-Bratzler shear force determinations.

Warner-Bratzler shear blade specifications include: (1) blade thickness of 1.1684 mm (0.046 inches); (2) V-notched (60° angle) cutting blade; (3) cutting edge beveled to a half-round; (4) corner of V rounded to a quarter-round of a 2.363-mm diameter circle; (5) spacers providing a gap for the cutting blade to slide through of 2.0828-mm thickness. After cooking and recording the final cooked temperature and weight, steaks, chops, and filets should be chilled overnight at 2 to 5°C before coring. Chilling firms the steak, chop, and filet and makes it easier to obtain uniform diameter cores or pieces. If chilling is not used, some protocol to obtain a consistent steak temperature before coring should be followed, such as allowing steaks to reach room temperature (23°C). Round cores should be uniformly 1.27 cm (0.5 inches) in diameter and removed parallel to the longitudinal orientation of the muscle fibers so that the shearing action is perpendicular to the longitudinal orientation of the muscle fibers (Figure 32).

Cores can be obtained using a handheld coring device (cork borer) or an automated coring device (drill press with cork borer attached). Coring devices must be in good condition and sharp, or the core diameters will not be consistent and will result in spurious increased variation in shear values. Hand coring requires very careful attention to the amount of pressure used as the corer is turned. It is best practice for the same person to do all of the coring throughout the duration of a study. A minimum of 6 cores should be obtained from each sample (this might require 2 pork chops or 3 lamb chops). Cores that are not uniform in diameter, have obvious connective tissue defects, or otherwise would not be representative of the sample should be discarded. If cooked steaks/chops were chilled, cores should be kept refrigerated and covered to reduce moisture loss until sheared to maintain a consistent temperature. Obtaining uniform cores is the most critical variable in WBSF measurements.

For poultry, uniform cores are difficult to obtain. Lyon and Lyon (1991) used 1.9-cm X 1.9-cm strip sections that run parallel to the fiber direction in the breast for shear determination. A minimum of 2 strips per filet are sheared perpendicularly to the fiber direction. Another sampling method, referred to as a strip technique (Schilling et al., 2003; Figure 33), is used by some labs. This technique uses paired scalpels that are bolted together so that the cutting edge is exactly 1.0-cm apart (G-R Manufacturing, Manhattan, KS). The paired scalpels are used to make a cut parallel to the muscle fibers. The paired scalpels, in essence, act as a subsequent cutting guide. Typically, 2 to 3 cuts are made with the scalpel using the previous cut as the guide for the next cut. After the initial cut, a sharp single-blade meat-cutting knife is used to complete the cut. The 1-cm thick strip is then tipped on its side to identify the muscle-fiber orientation for the next cut. Once the final cut is made, the width and thickness of the strip should be measured to confirm the 1-cm dimension has been achieved. One advantage of the strip technique is that many more samples for shearing can be obtained in comparison to the coring method. This particular method is very useful for obtaining samples from cooked broiler breasts.

Figure 33.
Figure 33.

Strip technique for poultry using 190 × 190 mm strip. Picture is an example of a Warner-Bratzler setup.

All values obtained should be used for mean calculation, unless visual observation indicates a valid reason a value should be discarded (e.g., a piece of connective tissue in the shear plane). Each core should be sheared once in the center to avoid the hardening that occurs toward the outside cooked edge of the sample. WBSF tests using automated testing machines should be conducted with a crosshead speed of 200 to 250 mm/min. Shear tests that do not follow these equipment or sample specifications should not be referred to as WBSF (such as square holes in the shear blade, square meat samples, straight-edged shear blade, or blade not properly beveled, etc.). For a demonstration of WBSF, watch the video at http://www.meatscience.org/sensory.

For tenderness of whole-muscle cuts, a critical problem has been the diversity in shear force measurement protocols that has made it nearly impossible to appropriately compare shear force values among data published by different institutions. The large variability in mean shear force and repeatability of shear force from different institutions and different protocols has been demonstrated by comparing WBSF and slice shear force (SSF) values on matched steaks from the same animals (Wheeler et al., 1997b; Wheeler et al., 2007). Those data highlight not only the importance of standardized protocols, but also limitations associated with applying tenderness thresholds, such as those reported by Shackelford et al. (1991) and Martinez et al. (2023), which relate WBSF values to trained sensory panel tenderness ratings. Because WBSF and SSF measurements differ substantially and are sensitive to protocol-specific factors, such thresholds may be valid only for data generated using the same protocol, within the same institution, or within a single study.

It is often useful to be able to compare objective measures among institutions, but also make useful comparisons of consumer tenderness/texture data with objective measures of tenderness/texture. Toward this goal, much progress has been made on the use of more standardized protocols. When there is not one methodology that is superior to others, the advantages of everyone having comparable data outweigh the disadvantages of a restrictive protocol. A superior method, however, should not be discarded for the sake of uniformity of methodology. It is also recognized that the recommendation of certain procedures in these guidelines should not discourage research to evaluate alternative methods or instruments.

Numerous devices have been tested for their ability to measure meat tenderness. The measurement most often used has been WBSF, which is still frequently used today. Beef longissimus WBSF of 1.27-cm diameter cores are highly repeatable when measurement protocols are executed properly (Wheeler et al., 1994. 1996, 1997a). Several sources of error have been identified that contribute to errors in shear force assessment within and among institutions (Wheeler et al., 1994, 1996, 1997a). Available data have not demonstrated that more accurate or repeatable data are obtained when WBSF is conducted on chilled samples compared to room-temperature samples (Wheeler et al., 1994). Thus, although empirical observation indicates it is easier to obtain more uniform diameter cores from chilled steaks than room-temperature steaks, steak temperature at coring may not be critical as long as it is consistent within an experiment.

Slice shear force

Shackelford et al. (1999a) developed a simplified technique for measuring beef longissimus shear force, which is referred to as SSF. The repeatability of SSF (0.89) exceeded repeatability estimates (0.53 to 0.86) that have been reported for longissimus WBSF (Wheeler et al., 1996, 1997b), and Shackelford et al. (1999a) reported that SSF was more accurate than WBSF. Because of time constraints associated with online assessment of meat tenderness, some aspects of the SSF protocol that Shackelford et al. (1999a) developed for online assessment of beef longissimus tenderness may not be necessary or desirable for routine collection of shear force data in a laboratory setting. Thus, Shackelford et al. (1999b) conducted a series of experiments to develop an optimal protocol for routine SSF measurement in research and to evaluate SSF as an objective method of assessing beef longissimus tenderness. Subsequent experiments were conducted to adapt this technique for use on pork longissimus, lamb longissimus, and many other beef muscles (Shackelford et al., 2004a and 2004b).

Because the SSF technique allows for a substantial increase in laboratory throughput, it has allowed for the development of novel tenderness management systems. To date, by far, the most common use of the SSF technique has been for the evaluation of beef longissimus tenderness. This has included routine testing of commercial product lines as well as in research. To conduct SSF, immediately after cooking, a 1-cm thick, 5-cm long slice is removed from each steak parallel to the muscle fibers for SSF determinations. The slice is acquired by first cutting across the width of the longest point, approximately 2 cm from the lateral end of the muscle. Using a sample sizer, a cut is made across the longissimus parallel to the first cut at a distance of 5 cm from the first cut. Using a knife that consists of 2 parallel blades spaced 1-cm apart, 2 parallel cuts are simultaneously made through the length of the 5-cm long steak portion at a 45° angle to the long axis of the longissimus and parallel with the muscle fibers.

The 5-cm long, 1-cm thick slice is sheared perpendicular to the muscle fibers using a universal testing machine equipped with a flat, blunt-end blade (Figure 34). The SSF blade is designed to replace the WBSF blade on an automated testing machine. The SSF blade has the same thickness (1.1684 mm) and degree of bevel (half-round) on the shearing edge as WBSF blades and should be used with the same amount of gap (2.0828 mm) for the blade to pass through during shearing. The crosshead speed is set at 500 mm/min to minimize the time required for measurement of shear force. Optionally, SSF could be measured using a WBSF machine equipped with an SSF blade, as described in the “small volume” protocol listed below. In that case, the crosshead speed is dictated by the WBSF machine.

Figure 34.
Figure 34.

Slice shear force blade segmenting a meat slice.

Recent data indicate SSF measurement on longissimus and gluteus medius can be conducted with slices obtained perpendicular to the steak/chop cut surface Evidence indicates SSF is much less sensitive to muscle-fiber orientation effects than WBSF.

The limited data available comparing hot vs. cold SSF, however, indicates hot SSF resulted in stronger correlations to WBSF and trained sensory tenderness ratings than cold SSF (Shackelford et al., 1999b). Thus, the implication of more accurate data and the increased convenience of cooking and shearing in the same day leads to a recommendation for hot SSF.

It has been clearly shown that shear force (both WBSF and SSF) does not properly reflect tenderness differences among muscles (Bouton et al., 1978; Harris and Shorthose, 1988; Shackelford et al., 1995; et al., 2004; King et al., 2009). For example, beef biceps femoris and longissimus have similar WBSF values, but longissimus is more tender as assessed by a trained descriptive attribute panel (Shackelford et al., 1995; Rhee et al., 2004). Therefore, it is inappropriate to use shear force to compare tenderness differences among muscles. This likely will lead to misinterpretation of data and false conclusions. Sensory evaluation should be used to compare the tenderness of muscles. When it is necessary to measure shear force in multiple muscles in a given experiment, we urge the investigators not to compare (statistically or otherwise) shear force among muscles and to report the data in such a manner that does not encourage the reader to compare shear force among muscles. Furthermore, we urge the investigators to use a qualifying statement either as a footnote (legend entry) or as a part of the statistical analysis section of the document such as the following sentence from King et al. (2009): “Slice shear force data were analyzed independently for each muscle because measures of shear force do not accurately reflect tenderness differences between muscles (Shackelford et al., 1995; Rhee et al., 2004), because shear force does not accurately represent the contribution of connective tissue to muscle tenderness (Bouton et al., 1978; Harris and Shorthose, 1988).”

Allo-Kramer shear value

A popular alternative to the WBSF method is the Allo-Kramer shear (AKS) method, especially for poultry, fish filets, and further-processed meats (Figure 35). For poultry, the AKS value (kg of force per g of meat) can be determined on surface samples (40 × 20 × 7 mm), 25 mm diameter cores that are cored parallel to the muscle fiber, or strips that are cut parallel to the muscle fiber. Samples should be taken in at least duplicate, and a minimum of 30 breasts should be evaluated per treatment if the data will be published in a peer-reviewed journal. This could be 10 samples per treatment per replication for 3 replications. The instrument should be equipped with a 10-blade AKS compression cell, a 500-kg transducer load cell, a 50, 100, or 200 kg load cell, and a crosshead speed of 500 mm /min (Sams, 1990). Samples must be weighed before shearing. Shear values are calculated by dividing either the total kilogram shear force or total Newton shear force by the weight of the sample (kilograms of shear per gram of sample; or Newtons of shear per gram of sample), and total energy (Joules/g) is calculated as the area under the texturegram. Lyon and Lyon (1991) found that the AKS values had a significant correlation (R2 = 0.86) with the sensory texture values of chicken breast meat deboned at various times postmortem.

Figure 35.
Figure 35.

Allo-Kramer shear force method: a) Allo-Kramer blades and cell prior to shearing; b) inside of Allo-Kramer shear cell when chicken breast sample (40 × 20 × 7 mm) oriented so that sample is sheared across the muscle fibers.

For further-processed meats, the AKS force is conducted similarly to the description for chicken breast and fish. However, samples will be taken such that they are a specific size. For example, a 25.4 mm diameter core that is 12.7 mm in length is an excellent sample size for this application. A 20 mm cube or sample that is 20 mm in width × 40 mm in length × 7 mm in depth is an acceptable size for use for AKS force of deli meats. The instrument should be equipped with a 10-blade AKS compression cell, a 500-kg transducer load cell, a 50, 100, or 200 kg load cell, and a crosshead speed of 500 mm /min. Shear values are calculated by dividing either the total kilogram shear force or total Newton shear force by the weight of the sample (kilograms of shear per gram of sample; or Newtons of shear per gram of sample), and total energy (Joules/g) is calculated as the area under the texturegram. Both shear force and total energy estimate protein binding strength.

Meullenet-Owens razor shear

Common texture analysis methods, such as AKS, WBSF, and texture-profile analysis (TPA), can be subject to variation due to excising the sample from the whole muscle or variation in texture due to location (Owens and Meullenet, 2010). Hence, there was a need to develop novel methods that can be used on whole samples, especially for poultry, thus reducing labor, time for analysis, and variability due to location. The Meullenet-Owens razor shear (MORS) method was initially developed for poultry meat (Cavitt et al., 2004, 2005a, 2005b; Figure 36). The purpose was to be able to shear intact filets rather than to have to cut samples to an exact dimension. Some advantages of this method are: 1) reduction of experimental error due to sample preparation; 2) the method is independent of sample dimensions; and 3) there is no sample preparation time (Lee et al., 2008). It is important to change the razor blade periodically to prevent dulling of the blade. Recommendations are to change the blade after 100 shear measurements, but it is important to be consistent within a study and make sure that the blade being used is sharp. A blunt version of MORS was developed, which has been used for tougher meat samples (BMORS; Figure 36). Lee et al. (2008) reported that BMORS was more discriminative than MORS at higher shear values, while MORS may be more discriminative for tender meat. With the blunt version, the shear test includes compression as well as shear. This results in slightly higher shear values (energy and force) compared to MORS. The MORS and BMORS are very highly correlated (r = 0.99) with each other using either total energy or maximum force values. Both methods are strongly correlated (r < −0.82; r = −.87 to −0.90 for MORS/BMORS to tenderness intensity, respectively) to consumer sensory attributes such as overall impression, tenderness acceptance, tenderness intensity, and tenderness JAR (Lee et al., 2008). For MORS, a razor blade (8.9 mm wide) penetrates the cooked meat sample to a depth of 20 mm, perpendicular to the muscle fiber. Multiple shears are performed for each filet and averaged. For BMORS, the same blade is used, but the razor edge is cut off, leaving a blunt edge. Replacement of the BMORS blade during testing is not necessary. All other parameters are the same. Cooked meat samples should be used for the assessment of tenderness properties when MORS and BMORS are used.

Figure 36.
Figure 36.

The Meullenet-Owens razor shear and Blunt Meullenet-Owens razor shear texture evaluation methods. a) MORS (razor edge) and BMORS (blunt edge) blades with close-up view of edges; b) MORS method with chicken breast filet; c) close-up of shear process as blade shears across muscle fibers. Abbreviations: BMORS, Blunt Meullenet-Owens Razor Shear; MORS, Meullenet-Owens razor shear.

For the MORS and BMORS test, a 20-mm-thick (minimum) piece of meat is placed on a flat aluminum plate with a 1 × 50 mm slot for the blunt blade (0.4 mm wide) to pass through. The texture of the sample is assessed using the BMORS or MORS device on a texture analyzer (TA-YT2i, Stable Micro System, Ltd). While samples less than 20 mm are not recommended, there may be times when larger samples are not possible. In these cases, it is recommended to report the maximum force (N) rather than the total energy. This system has been used with chicken breasts but has also been used for deli meats as described by Luckett et al. (2014). Using the data acquisition software (Texture Exponent 32, Stable MicroSystems, Ltd), 6 parameters are defined: positive force for springiness, positive area for hardness, time to positive peak for cohesiveness of mass, negative force for fibrousness between teeth, negative peak for rubberiness, and gradient to positive peak are obtained. The crosshead speed should be fixed at 5 mm/s, and the sample shear depth at 20 mm with a trigger force of 10 g. A multiblade shear test was developed by Zhang et al. (2021), which showed efficacy when evaluating broiler breast meat with or without myopathies.

Texture-profile analysis methodology: 2-cycle compression

Texture assessment of further-processed meats can be assessed using the texture assessment methods defined above. However, assessment of physical properties in further-processed meats is needed to understand texture. A common method of texture assessment is the TPA method. This method entails 2 sequential compressions on the same sample (Bourne, 1978; Montejano et al., 1985; Claus, 1989) and is classified as a dual-compression test. Measurements obtained using TPA that are commonly applicable to processed meat type samples include: hardness, cohesiveness, springiness, and chewiness (Table 24). Less commonly measured is adhesiveness. Hardness is measured with a single compression test, and the other 4 measurements are from the dual-compression test. Proper setup for performing TPA entails some pretesting of the samples that will be analyzed so that the instrument is correctly configured. Using a capacity load cell that is too high a may not be sensitive enough for delicate samples. With a too-low-capacity load cell, some samples may exceed its capacity. The following are key elements/steps that should be observed.

Table 24.

Texture-profile analysis traits and their description

Trait Description
Hardness Peak force at first compression.
Cohesiveness Integrated area under the curve of the second compression (A2) divided by the integrated area under the curve of the first compression (A1).
Springiness Represents the height that the sample springs back prior to the start of the second compression.
Chewiness This is a product of hardness times cohesiveness times springiness.
Adhesiveness This is the negative force that occurs when the compression platen is pulled away from the sample after the downstroke.

A standard configuration sample should be predetermined. In general, the shape of the sample should be such that it retains its basic orientation during compression. For core samples removed from a sausage, the diameter of the core should be at least equal to the height of the sample, but preferably a little larger. The core should be removed perpendicular to the cut surface. For sausages that are not to be cored, the sections should be cut out perpendicular to the long axis, again with the samples being wider than they are tall. If the intent is to perform TPA on non-cored samples so that the contribution of the skin (coagulated protein) can be included, this can be done if the sausages have a fairly equal diameter among the products being evaluated. If not, then cores should be used, and an alternative test for the skin contribution (e.g., puncture test) should be used. Samples should be presented to the TPA by being placed on their flat cut surface. Important aspects of evaluation include: 1) the temperature of the samples at the time of compression needs to be standardized and documented; 2) the compression platen (plate) that is selected must be large enough so that when the sample is compressed to the selected final compression percentage, none of the sample is outside the outer edge of the plate; 3) ideally select a load cell in which the maximum peak force of the sample will fall in the middle of its capacity; 4) the texture analyzer should be calibrated before testing; and 5) set the height of the compression platen (plate) adequately above the sample so there is room to place and remove the sample as well as inspect the surface of the compression plate for residual sample post compression. TPA instruments vary in parameter settings and therefore need to be specified. Settings that should be documented include: 1) pre-test, test and post-test speed, usually set at 5 mm/s; 2) trigger force is the threshold for the instrument to recognize contact with the sample; 3) fracture detection parameters; 4) delay time between compressions; 5) percentage compression usually set at 75%, 70% or 60% of sample height.

Careful attention should be given to how the instrument designates measurement criteria. The word descriptions of the TPA traits help provide a general understanding of the physical measurements. However, in the literature, there are a few variations on how some of the parameters are actually calculated. The more traditional method of calculating the major traits is detailed first (Table 24).

Figure 37 shows a typical 2-cycle TPA compression curve in Newtons. Hardness, which is the peak force (N) of the first compression, relates quite well to sensory firmness. It is important that this first compression does not break the product, since it will prevent the accurate measurement of other items. Cohesiveness is determined by independently calculating the integrated area under each compression curve. The area under the first compression curve (A1) starts from the initial contact (point P1) with the sample up to the completion of the designated percentage compression (P2). Note that it was not defined as the point at which the maximum peak force was achieved. The reason for this will be made apparent when discussing a sample that has a significant fracture point. The integrated area under the second compression (A2) starts when the compression plate comes in contact with the sample (P3) and concludes when the same downstroke distance of the first compression has been achieved (P4). Cohesiveness is then represented as a ratio A2/A1 (no units). In some cases, cohesiveness is presented as a percentage (A2/A1 * 100).

Figure 37.
Figure 37.

Example of a texture-profile analysis 2-cycle compression curve. Abbreviation: TPA, texture-profile analysis.

Springiness is determined as the base width of the second compression (b2 = P4−P3, mm). Since comparisons are often made between products, using only b2 for springiness is sound if the exact sample height is controlled. To account for some variation in sample height, springiness can be calculated as b2/b1, which then puts it on a ratio basis rather than an absolute recovery height measured basis. In some cases, springiness is presented as a percentage (b2/b1 * 100).

Chewiness is the product of hardness times cohesiveness times springiness (N × mm). In many respects, the percentage compression used to determine hardness should be greater than that used to determine cohesiveness and springiness. This means a sample from the same product should be compressed more extensively (e.g., 75% of the sample height), and then a second sample should not be compressed as much (e.g., 60%). For hardness, the sample should be more fully compressed so that it is indicative of what a sensory panelist evaluates for firmness. If the sample is over-compressed, there may be so much structural damage from the first compression that the second compression may only show a minor compression curve. This is why pretesting samples is of value to determine an applicable compression percentage. Nevertheless, if a direct comparison needs to be made to a published paper, the same percentage compressions should be utilized.

Another example of a 2-cycle compression curve is illustrated in Figure 38. In this example, the sample peak during the first compression did not correspond to the end of the downstroke (specified percentage compression of the sample height). The end of the compression downstroke occurred at P2. Users are cautioned to be sure that they know where the completion of the downstroke occurs for integrating the area under the curve.

Figure 38.
Figure 38.

Two-cycle compression test with 2 peaks during the first compression. Abbreviation: TPA, texture-profile analysis.

An alternative calculation for TPA cohesiveness that is sometimes used is to integrate the total energy under the curve defined as A1 and A2 (Figure 39). In this method, cohesiveness is calculated as A2/A1. Note that the endpoints of P2 and P4 represent points where the curve goes to zero or levels off after the compression cycle has completed. This approach avoids the need to validate the completed downstroke position as discussed in Figures 37 and 38.

Figure 39.
Figure 39.

Example for calculating alternate cohesiveness using texture-profile analysis. Abbreviation: TPA, texture-profile analysis.

TPA users need to fully disclose how they calculated all traits, and in particular, values for chewiness will be difficult to interpret without knowing if cohesiveness was calculated as a ratio or percentage, or if springiness was only based on the width of the second compression, a ratio (b2/b1), or a percentage (b2/b1 * 100). For ground beef patties, TPA can be conducted using one core (2.54-cm diameter) removed from the center of 10 cooked patties and compressed twice to 70% its original height to determine hardness (N), cohesiveness (ratio), springiness (height sample recovers, mm), gumminess (product of hardness times cohesiveness, N), and chewiness (N x mm). A crosshead speed of 100 mm/min should be used. AKS force can also be used. Using AKS, 1 strip (2.5-cm wide) should be cut from the center of each of 10 patties per treatment and replication. Each strip should be analyzed by either shearing one with a multiblade Allo-Kramer shearing device or shearing 3 times with a single-blade, straight-edge Allo-Kramer shearing device. Texture has also been assessed where these strips are sheared 3 times with a straight-edge SSF blade attachment. In each situation, a crosshead speed of 200 to 250 mm/min should be used for an automated testing machine.

Texture-profile analysis methodology: 1-cycle compression

One-cycle compression tests can be used to understand the underlying structural component of meat products. While it is used similarly to 2-cycle compression tests, its use is less common, and there is limited information compared to the 2-cycle TPA test. The test should be conducted as defined for the 2-cycle TPA test; however, only one compression is used. Data collected are peak force, A1, B1, P1, and P2 as defined in Figure 37. Not all attributes can be calculated in the one-cycle test as defined for the 2-cycle test.

Protein binding strength can be determined using a procedure that is modified from Field et al. (1984), as illustrated in Figure 40a. For this application, the diameter of the deli meat or ground meat patty should be between 10 and 12.5 cm in diameter. However, a machine shop could make an attachment specific to fit texture analysis instrumentation that could be used for meat pieces with other dimensions. This device was designed for 12.7-mm-thick deli slices, but could be used for products with other thicknesses. The thickness of the slice should be between 3.2- and 12.7-mm thick and should be the same for all samples. A 25.0-mm diameter steel ball (chrome alloy grade 25) is attached to a rod that is attached to the texture assessment instrument. Nails are placed manually through each sample into 1.6-mm holes drilled on the top of a plexiglass stand to secure the deli meat slices during testing. A minimum of 4 slices should be evaluated from each treatment for each replication due to biological variability in meat products. Nail holes were drilled 0.5 mm apart in 1 mm deep holes around a circle with a nail diameter of 4.5 mm and an inside diameter of 4.0 mm. The plexiglass stand is placed on the flat, circular surface of the texture analysis instrument. The slice should be aligned so that the steel ball penetrates the middle of the meat slice. The steel ball is positioned directly above the meat slice, and protein bind is estimated as the peak force (N) for the polished steel ball to burst through the slice of deli meat at a speed of 100 mm/min. If the meat slice breaks in a place other than the center of the slice, that data should not be used (Daigle et al., 2005).

Figure 40.
Figure 40.

Ball head (a) and star probe (b) used in compression tests.

A modification of the penetrating probe/steel ball test was developed to determine if there was enough functional protein in a processed meat product. For example, this test can be used to determine if mechanically separated poultry meat has sufficient protein binding in Newtons (N) to form a frankfurter. This testing was developed to evaluate if mechanically separated chicken (MSC) that was treated with antimicrobials to control Salmonella was not denatured by the antimicrobial dips and sprays that were used. To conduct this test, MSC is weighed and preblended with 1% NaCl. For each treatment, six 11.5 cm × 11.5 cm patties that weighed approximately 180 ± 0.1g should be formed and baked to an internal temperature of 74°C on a broiler pan with a rack (Viking Professional, Greenwood, MS, USA) at 204°C. Similar instrumentation could be used for other-sized patties, but these sizes work well with the equipment that is available. After reaching a temperature of 74°C, samples are removed from the oven and cooled to room temperature (22 ± 2°C). The instrument that is used to measure texture should be equipped with a 500 N load cell and a 25.0-mm diameter steel ball (chrome alloy grade 25). For each sample analysis, a patty is centered on the plexiglass stand so that the steel ball penetrates the middle of the sample once. Compressive load at maximum peak force (N) estimates protein binding value. Values greater than 20 N generally have sufficient protein to bind to form a frankfurter with acceptable water-holding capacity and texture.

Star probe

The star probe was developed as an alternative to WBSF and TPA evaluation for pork and further-processed meats (Figure 40b) and has a 5-pointed cutting edge. The sharpness of the blade edges should be maintained. Crosshead speed is up to the user and should be defined. For all products, the sample needs to be sufficiently thick so that the star probe can fully penetrate the surface without contacting the base that the sample rests on. Multiple probe measurements can be made on the same sample, provided that there is sufficient space between each penetration. For intact muscle cuts, a minimum of 2.5 cm between penetrations should be used. The star probe can be used for finely comminuted sausages. Investigators need to determine the minimum distance between penetrations that does not alter the structural integrity of the subsequent penetration. The star probe should not be used on ground patties as they tend to be more fragile, and the data will not provide a true measure of penetration.

Tensile testing

Tensile testing is a method that provides information on the strength of the bind within a further-processed meat product (Honikel, 1998). To conduct this method, one tensile grip should be attached to the base of the instrument that is used to analyze texture, while the other grip should be attached to the load cell. Initial grip separation can be variable, but 28.0 mm has been recommended by Honikel (1998). Crosshead speed should be approximately 1.0 mm/s (50–60 mm/min until rupture). This is based on a strain rate of approximately 2 per min. Therefore, if the strip is 28 mm in length, an extension rate of 56 mm/min is recommended. Honikel (1998) recommended a 28-mm-long strip that is cut out in the shape of a dumbbell, such that there is a 4 to 1 to 0.5 length to width to thickness ratio for the sample. For example, at 28 mm length, there should be a 7 mm diameter in the interior portion of the dumbbell (18 mm in length), and the entire strip should be 3.5 mm thick. The outer ends (5 mm long on each side) of the dumbbell should be approximately twice the width of the interior portion of the dumbbell. In addition, the sample should be continuously cut so that there is a smooth sample parallel fiber direction. The purpose of the dumbbell shape is such that the sample tears in the center of the strip and not on the ends. If the meat sample is not susceptible to tearing in the middle, the dumbbell shape is not necessary. For the test to be acceptable, the sample needs to break in the thinner portion of the dumbbell, in the center, such that it occurs in the parallel-sided region of the sample. A minimum of 8 samples should be evaluated per treatment, and only samples meeting the above criteria should be accepted as data. Breaking stress should be reported as N/m2 (Pa).

The following calculations are often reported: Breaking stress (N/m2 (Pa)) = (peak force)/(measured width x thickness) s; Breaking strain (extension of peak force/sample length) is the relative elongation of a material at its breaking point and is unitless.

Total energy to fracture (Joules) = area under the curve.

Blunt rod for skin assessment

The skin strength of further-processed meats can be affected by processing, cooking, ingredients, and other factors. Samples are prepared as defined for penetrating ball testing, except that a blunt rod is used to penetrate the skin of the sample. Samples can be evaluated cold or hot as long as a consistent temperature is maintained. Penetration is defined as when the peak force is reached, and the curve shows failure. Data are reported as peak force in N or kg. Blunt-end rods should be of sufficient circumference to provide reasonable measurement of the penetrating force, but not so large that the force is more indicative of compression. The diameter of the blunt-end rod is dependent upon how delicate the skin is on the processed meat and the capacity of the load cell installed on the texture analyzer.

Slice integrity

Slice Integrity is used for sliced deli meats to understand protein binding within the slice. For this method, each deli product treatment is sliced into 100 1.6 mm slices. Slice integrity is determined by the number of slices out of 100 that remain intact (no holes or tears) after slicing and holding in the air by fingertips to ensure that it does not tear before packaging. Slice integrity is an indicator of protein binding.

Torsion testing

Torsion testing is a useful texture analysis method for use on finely comminuted sausages and studies on the fundamental properties of protein gel systems. As the name implies, the sample is mounted in a device that applies torsional force as a result of twisting the sample until the point of failure (Figure 41a). This method yields 2 independent mechanical properties, which are stress (kPa) and strain. Stress is the peak force that occurs at the point of failure. Stress is sometimes also referred to as rigidity or fracture stress. Strain is unitless but can be thought of as the distance or time necessary to reach the failure point. Strain has also been described as deformation, that is, the extent to which the sample has to be deformed to reach its failure point. Cooked sausages or protein gels that are tested using a torsion gelometer must be homogeneous and isotropic. Isotropic is a key feature in that no matter how the sample was removed or obtained from a cooked sausage, the sample mounted on the torsion device must be independent of orientation. As such, finely comminuted processed meats adhere to this requirement. Frankfurters and bologna-type sausages are compatible with this method. In the case of frankfurters, a standard section from the sausage is cut perpendicular to the length (Figure 41b), and then Mylar discs with alignment notches are glued (cyanoacrylate) on each end. The sample is then milled (Figure 41c, Accu-Tool, LLC, Apex, NC, USA) to an hourglass shape to a standard diameter (10 mm). This, of course, removes the coagulated skin layer from the sausage and assures that the failure will occur independent of the device that holds one end stationary while the other end is twisted. Basically, any finely comminuted sausage or protein gel that the milled hourglass-shaped sample can support itself (stand erect; Figure 41d) will enable torsion testing.

Figure 41.
Figure 41.

Torsion testing with comminuted meat products (a) where the whole system is set up, (b) cutting and gluing a frankfurter section, (c) a core being milled using a Accu-Tool, (d) to an hourglass shape (d) to a standard diameter (10 mm).

Spurious results may be obtained with this texture method if there are significant air pockets during the sausage manufacturing process. If an air pocket is present in the narrow portion of the dumbbell-shaped sample, it will prematurely fail (tear). Obviously, samples with visual air pockets on the outside of the milled sample should be excluded from use. However, after the sample has been processed on the torsion device, the interior of the sheared surface should be assessed for flaws.

Instrument calibration

Calibration is essential with any instrument, and verification is recommended at 12- to 18-mo intervals. Calibration is the daily spot-check of the instrument’s accuracy. It is performed by placing a known weight on the transducer or by applying a known voltage to the cell via a shunt. Adjustments can be made so that the instrument output matches the known weight or voltage input. The calibration procedure is usually performed using a single value or with weights at approximately 20% and 80% of the force values expected in the test (see ASTM Standard E4-24, 2010, for other approaches). Some manufacturers recommend that the shear blade attachment be in place during calibration so that its weight and any side drag are accounted for during taring. Many protocols calibrate and then tare the load cell. The calibration procedure should be defined and followed during testing. Automated testing machines should be verified according to the manufacturer’s instructions. When crosshead speed is critical to the results, it is recommended to have it verified as well.

When force values form the basis of value, for example, premiums for tenderness, instruments will need to be validated daily (ASTM Standard F2343-06, 2006). This process is referred to as validation, a procedure that documents, confirms, and provides assurances that force measurements will consistently meet predetermined specifications and force attributes. Further, validation also provides assurances that instruments are installed and operated within the manufacturer’s specifications (ASTM Standard F2341-05, 2005).

The USDA’s Livestock, Poultry and Seed Program approves third-party laboratories for conducting SSF/WBSF testing to ensure that the requirements of an ASTM International tenderness standard have been met. Third-party laboratories are approved to conduct SSF and/or WBSF through proficiency testing conducted in conjunction with an AMS-designated SSF/WBSF reference laboratory (USDA Livestock and Poultry Program, 2012).

Data Analyses

Data analysis is a critical component of sensory evaluation. Sensory data are, in the true sense, multivariate. When a human, either trained or a consumer, evaluates a meat sample, they utilize multiple senses during the evaluation process. Trained, descriptive attribute sensory panelists are educated to incur sensory input, segment the information into single attributes, and then rate the intensity of each attribute using an anchored, defined scale. Consumers incur sensory input and may or may not segment that information into individual attributes, unless specifically asked to rate a specific trait. Therefore, overall liking/disliking consumer ratings are really ratings of multivariate information, and consumer ratings of specific attributes—such as liking or disliking of overall flavor, overall texture, overall juiciness, intensity of beef flavor, or level of juiciness—are univariate attributes.

There are multiple statistical analysis tools for sensory data. Researchers should focus on tools that most effectively address the hypothesis of the study. This section will provide information on the most commonly used data analysis tools. Univariate and multivariate tools will be discussed for descriptive and consumer data. Additional information can be obtained in Meilgaard et al. (2025), O’Mahony (1986), and ASTM Committee E18 documents on statistical analyses.

A statistician or sensory professional with extensive experience in analyzing data should be consulted before conducting any sensory experiment. The experimental design, randomization, and hypothesis should be clearly defined before initiation of data collection. Initial statistical models and level of significance should be defined so that the experiment is properly executed and predetermined biases do not inadvertently affect the experimental results. Statistical power test calculations should be used to determine replicate numbers based on conditions defined above.

Data preparation

Sensory data are traditionally either collected using a handwritten ballot filled out by each panelist or collected into a computer software program. Regardless of the collection procedure for the data, it needs to be organized so that one row represents the responses of one panelist for one sample. Variables for sensory day, session within sensory day, order, 3-digit sensory codes, panelists, treatment codes, and sensory variables for each response should be included as columns. Variables that might be potential covariates or blocks, such as cook time and cook yield, should be included.

If data have been entered into a spreadsheet with panelists’ results for each sample reported, then reorganization is not necessary, and the next step of analysis can begin. If data have been collected using a computer program, the data might need to be unscrambled from the randomized sample sheet to assign the observed values to the actual treatments using a program like Unscrambler from SAS or within Excel (Microsoft). If handwritten ballots are used, data should be entered into a spreadsheet and then verified by independent personnel using the original data sheets. Computerized data sheets should be checked for layout and format. If a panelist is not present for a sensory day or sample, the missing data should be indicated by a “.” or the cell left blank, depending on the statistical analysis program being used. If a panelist evaluates an attribute and does not find it present, a numerical value of “0” should be entered into the cell if 0 = none or absence of the attribute. Summary statistics for mean, standard deviation, standard error, minimum, and maximum values should be calculated either within the spreadsheet or by a data analysis program. Summary statistics can be useful in conducting diagnostics to ensure that data are within range and that gross errors in data entry are found.

It is assumed that researchers have experience using the statistical package for data analysis of their choice. Many statistical packages are available for use, such as SAS, SPSS, R, JMP, XLSTAT, and others. Resources to understand programming language should be used to implement the data analyses, and will not be covered here. If a data program, like Compusense, is used, it is the responsibility of the sensory professional to understand how the data is being analyzed and to verify that the model and experimental design match.

After assurance that the data are correctly entered, tests for normality are conducted. Consumer data are commonly abnormal in distribution, but all data should be tested for normality because normality is an assumption when conducting ANOVA. Data can be analyzed with non-normal assumption procedures like PROC GLIMMIX of SAS that generalize the MIXED procedure. The GLIMMIX procedure generalizes the MIXED and GENMOD procedures in 2 important ways. First, the response can have a non-normal distribution. The MIXED procedure assumes that the response is normally (Gaussian) distributed. Second, the GLIMMIX procedure incorporates random effects in the model and thus allows for subject-specific (conditional) and population-averaged (marginal) inference. If data are not normally distributed and the data are analyzed using a model that does not account for non-normal distributions, the data should be tested before analysis. Normality of data can be tested in many ways. Some statisticians recommend plotting frequency distributions of the data and examining the shape of the curve. If the curve is bell-shaped, the data are normally distributed. This method, however, is not highly accurate. Most statistical programs have methods to test for normality. In SAS, the Proc Univariate function performs many descriptive statistics for normality, such as Q-Q plot, stem-and-leaf plot, box plot, and normal-probability plot. This procedure also conducts Kolmogorov-Smirnov, Shapiro-Wilk, Anderson-Darling, and Cramer-von Mises tests to evaluate normality. The Box-Cox transformations in SAS can be used to determine logarithmic transformation constants that can be used to convert the data to a normal distribution. Other transformations also might be appropriate. If data are transformed—for ease of interpretation—least-squares or unadjusted means should be transformed back to the original scale, but least significant difference values from the original scale cannot be used. The researcher should also check to ensure that the back-transformed means are not substantially lower than the means in the original data. This can be an issue when the transformation is logarithmic.

A second way to test for normality is to plot the residuals of the data. It is suggested, however, that if there are missing observations, a “means” statement may not be used, and the researcher should use least-squares means to compare treatment means. In sensory data, it is not uncommon to have a panelist missing from a day of evaluation; in that case, least-squares means should be used. Diagnostics are run to ensure that data comply with the assumptions of an ANOVA; i.e., model errors are normally and independently distributed random variables with mean = 0 and variance = σ2, constant variance and homogeneously dispersed, and the measurements were acquired in a random order in a completely randomized design (Huntsberger and Billingsley, 1979; Steel and Torrie, 1980; Devore and Peck, 2005). To run diagnostics, the predicted and studentized residual values of the data need to be calculated and obtained from the ANOVA procedure. The predicted and studentized residual values can be plotted as in Figure 42. For example, the “gplot” procedure in SAS will plot the studentized residuals by the predicted observation values to determine whether or not the variance is constant and homogeneously dispersed. If the variance is constant and homogeneously dispersed, then the plot should look similar to the scatter plot in Figure 42, where there is no pattern to the residuals. Normality of the data can be assessed using the “gchart” procedure, and if the data are normally distributed, then the histogram should somewhat resemble a bell-shaped curve. It should be noted that the ANOVA is not very sensitive to departures from normality or unequal variance and is a relatively robust statistical test (Ott and Longnecker, 2001). Slight deviations from normality will most likely not affect analyses and interpretation of results. If there is concern, a statistician should be consulted.

Figure 42.
Figure 42.

Example plot of residuals (a) to evaluate whether variance is constant and homogeneously dispersed and (b) histogram to evaluate normality of the data.

Significance levels for analysis should be predetermined. While an α-level of < 0.05 is commonly used, in a small dataset, an α of < 0.10 could be justified, whereas in a large dataset, an α of < 0.01 might be more appropriate. It is important to predetermine significance levels based on experimental design and observation values before initiation of the study to prevent potential biases when interpreting sensory data. In addition, α-error should always reflect the level of confidence that the researchers will require to make inferences before any data are collected or analyzed.

Descriptive sensory data analysis

Most descriptive sensory data that is generated using trained panels are analyzed using ANOVA, where each sensory attribute defined in the ballot is evaluated for the effects of treatment independently. This type of analysis provides the opportunity to understand if specific attributes differ among treatment levels. For trained descriptive data, an initial model that includes random (sensory day, session, and order) and fixed (treatment and subsequent interactions) effects should be developed to reduce the risk of encountering β-error. Potential covariates should be included in the model if they help to explain the error. Random variables that account for additional variation, and to make inferences about the population from which the experimental units were derived, can be included. In meat science data, slaughter day, processing day, shift of production, and plant where experimental units were collected are examples of random variables. Meat scientists many times conduct repeated-measure experiments, especially when conducting studies that include different storage times or when multiple steaks are removed from a subprimal and then randomly assigned to treatment. It is wise to consult with a qualified statistician to ensure that the analysis is being conducted correctly and that the error terms that are used to test effects are appropriate.

Panelist effects

With trained descriptive data, the researchers have the unique opportunity to understand the efficacy of their panelists. In the statistical model, panelists would be considered a pseudo-replicate. Extensive training, performance evaluation, and validation of trained panelists are conducted before conducting a study, as previously discussed. If these procedures are followed, the sensory professional has assurance that panelists are consistently and uniformly evaluating samples. During testing, however, factors can influence sensory verdicts, and data should be evaluated for panelist and treatment by panelist interaction effects before final analyses. It should be understood that the more training and experience a panel has with a product, the more sensitive the panel. To test panelist and panelist by treatment interaction effects, an experimental unit is defined as an individual panelist’s response to a sensory sample. For example, 6 panelists independently, in separate booths, evaluated 54 steaks over 8 d, resulting in 324 observations. There were 6 treatments in the study with 9 replicates per treatment. To test panelist effects, a full model would be defined where panelist (n = 6), treatment (n = 6), and panelist by treatment (n = 36) would be fixed effects. Random effects of sensory day (n = 8) and order would be defined. In some studies, order would be included as a fixed effect if order was balanced (each treatment is included in each order equally across sensory days). Steak, in this case, or the experimental unit (n = 54) would be included in the model. The random error mean square error in the model is used as the error term to test the panelist and panelist by treatment effects. A significant panelist effect, using the predetermined level, indicates that panelists did not evaluate an attribute in the same way. If a significant interaction exists, differences can be meaningless or meaningful. It is up to the panel leader and researchers to understand these effects and their interpretation. When an interaction is significant, the first step is to reevaluate data accuracy. If a decimal point is improperly placed or a variable is incorrectly entered in a small study, an interaction could result. It should be noted that it is not unusual for panelist effects to be significant. This does not necessarily mean that the panelists or some of the panelists are not doing a good job in evaluating the products. If the panelists have been trained to detect a 1-point difference and the root mean square error is 0.50, then 0.5 differences might be significant. When examining the least-squares means between panelists, there might be only a 0.75 difference between the panelists scoring the lowest and highest within an attribute. In this case, the sensory panel leader would not be concerned about the panelist effect as panelists were expected to vary within 1 point. On the other hand, if, when examining the least-squares means, one panelist is scoring 2 to 3 points different from the remainder of the panelists, additional training is warranted.

If a significant panelist by treatment interaction is determined to be meaningful, the data should be presented as an interaction and not averaged across panelists. Theoretically, if the sensory panel leader has adequately tested the panelists’ readiness through performance evaluation, reference samples are provided for standardization, and panelists conduct warm-up samples daily before testing for calibration, panelist by treatment interactions should be minimized. If, during testing, a panelist rates samples within a treatment differently from other panelists, averaging across panelists might result in a misinterpretation of data. This also might be a result of variation within subsamples, and this information might be important in the interpretation of the data.

If a panelist is erratically evaluating treatments and is obviously different than other panelists, there might be sufficient justification for removing their data from the dataset. Careful consideration should be given to the decision to remove panelists because if the panel leader has sufficiently trained panelists, panelists have passed performance evaluation, and appropriate and consistent warm-up samples are used to calibrate panelists daily, it should not be automatically assumed that the panelist is rating samples inadequately. A panelist who is not evaluating the samples within the range or levels of other panelists, however, should be targeted for additional training. The panel leader needs to carefully evaluate this panelist’s performance and determine whether or not to keep their data.

Univariate analysis

If there is no concern for the efficacy of the panelists, the next steps in analysis can occur. There are 2 choices. Data can be averaged across panelists so that a mean value represents the intensity of a sensory attribute for each steak, or the original full model can be used, but only if the mean square error for steak or experimental unit is used to test the treatment effects, not the residual mean square error of the full model. In the example, if the data is averaged over panelists, there would be 54 lines of data where each steak would be represented by a line, and steaks are the replications. The model would include sensory day and order as random effects and treatment as a fixed effect. The residual error would be used to test the significance of the treatment. If the original full model is used, a statistical package where error effects can be defined to test effects would be used. To test the treatment effect, the mean square error for steak would be defined as the error. Note that in both scenarios, the error used to test the treatment effect is the random error associated with steak.

Examination of the full model ANOVA should be conducted. Interaction effects that are not significant can be removed (Meilgaard et al., 2025). There is no uniform consensus on when to remove or pool interaction effects into the error term. If the p-value for the interaction effect is large (greater than 0.50), it is obvious that the interaction is not partitioning variation from the model, and removing this interaction will not affect the outcome of the analysis. Depending on the sample size and the p-value for effects of interest (treatment and treatment interactions that are close to being significant), however, removing interactions that have P values between 0.05 and 0.25 can affect the level of significance of the effects of interest. When that occurs, the ANOVA tables with and without the interaction should be examined and compared. If removing interactions with P values greater than 0.5 affects the outcome, and they were not significant in the original model, the researchers must decide what to include in the final model. The conservative approach is to include interactions with P values less than 0.25 and exclude interactions with P values greater than 0.25. Researchers must understand how the significance of effects of interest is affected in both models for the correct interpretation of results. In studies with large numbers of observations, it is likely that removal of interactions that have P values greater than 0.25 will not appreciably affect the significance level of tests of hypotheses for effects of interest. Covariates that are not significant can be excluded from the full model; however, covariates that account for some variation (P values greater than 0.05 and less than 0.25) can be included in the model to explain error, but they are most likely not affecting the experimental outcome. After consideration of P values for all effects defined in the full model, a final model is defined, and the ANOVA is calculated. Least-squares means are usually calculated for trained sensory data because it is not uncommon to have missing or unbalanced subcell numbers.

If the final model includes the effects of panelist and panelist interactions as discussed above, treatment effects that might be included in the model but are not part of the panelists by treatment interaction may need to be tested using a different error term than the model residual error, as discussed above. If you are unsure of the proper error term for testing specific effects, consult a statistician. Many statistical packages allow for the definition of error terms for individual effects included in the model and may provide methods to determine the correct error term to use.

In the final model, significant main effects and interactions should be identified. Effects that are not significant can be as important as effects that are significant. If a significant interaction occurs upon analysis of multiple factors in a model (excluding block effects), then only the interaction effects should be presented and interpreted. For confirmation that the proper analysis was used, consult a statistician.

Smaller datasets

In meat science research, not all experimental designs include animal, carcass, primal, or subprimal experimental units, especially in processed meats, fish, poultry, and small protein sources. These are situations where multiple panelists cannot evaluate from the same protein unit, and multiple units are sampled from a replicate. For example, shrimp harvested from a day are defined as a batch, and 20 shrimp are randomly selected for sensory evaluation, or a meat blend is formulated and defined as a batch for sensory testing.

In these experiments, where a batch is defined as an experimental unit, limited experimental units are used. When analyzing the data, panelists and panelist by treatment effects should be examined. However, collapsing the data over panelists is most likely not possible due to the limited degrees of freedom in the full model. As panelists represent a source of variation, many times the panelist effect is maintained in the model, and the residual error may be used to test treatment effects. This model may not be as statistically powerful due to reduced degrees of freedom, but the model tests the treatment effects, and conclusions can be made. Another issue with these types of experiments is that variation within and between batches may be small, which results in difficulty in accounting for and partitioning variation in the statistical model. It is imperative to develop statistical models before testing and to discuss these models with an experienced professional to validate the model.

Treatment effects

After completion of the full model, determination of significant and nonsignificant effects, the least-squares means should be calculated. Nonsignificant and significant least-squares means should be reported with an estimate of error, either standard error or a root mean square error. If significant differences are reported for the effect, differences in least square means can be calculated using a mean separation technique. Statistical packages contain multiple mean separation techniques, and understanding the strengths and weaknesses of each technique is the responsibility of the sensory scientist.

Consumer data analyses

For consumer data, “Consumer” is included in the model since consumers normally account for the greatest source of variation. Consumers are untrained, and variation associated with consumers indicates differences in consumer perception or liking. Consumers can be evaluated using multivariate tools to understand potential segmentation or categorization of consumers that may help in explaining their source of variation (see discussion below).

When analyzing consumer data, it is also important to look at the percentage of “Like Extremely” and “Like Very Much” ratings (top 2 boxes in a 9-point hedonic scale) as well as “Dislike Extremely” and “Dislike Very Much” ratings (bottom 2 boxes) to check for shifts in rating patterns among treatments. It is possible for mean scores to not be significantly different but have shifts in top 2/bottom 2 box ratings, indicating consumer segmentation or polarized ratings.

Consumer data should be analyzed similarly to descriptive sensory data, where covariates, fixed and random effects, model considerations (repeated measures), and interaction effects are included in a full model. Consumer is included as an effect in the model and, in most cases, contains the highest source of variation. Location, if there are multiple locations, is included as fixed effects and is used to understand if there are differences in consumer sensory perceptions across locations. An example would be if consumers from the west, east, and southeast evaluated beef steaks. This partition in the model would assist the sensory professional in understanding if there were regional differences in consumer perception. Random effects of sensory day, session (if more than one session is conducted per day), and order are included in the full model. These effects account for variation in the model but are usually not used to interpret findings. The full model can be defined, and interaction effects may or may not be included in the model.

Data are presented as discussed for descriptive sensory data, but using consumer attributes. It should be noted that in all sensory data, least-squares means or means should be reported to 1 decimal place more sensitive than the original data (i.e., if the panelists are evaluating an attribute to the nearest 1 unit (1, 2, 3, …), data should be presented to the nearest 0.1 unit) and the measure of variance should be reported to 2 units more sensitive.

Multivariate analyses

Sensory data are by nature multivariate. When a human evaluates a sample, activated sensory receptors send information to the brain. Trained sensory panelists are proficient in either ignoring information (i.e., ignoring information that would be considered based on visual appearance and its relationship to flavor), segmenting or pulling out a specific attribute from the sensory stimulus (i.e., juiciness is evaluated independently of muscle-fiber tenderness or beef-flavor identity is measured independently of cooked beef fat), and quantifying it. Consumers might rate their overall liking of a sample and then be asked to evaluate specific attributes, such as their overall liking or intensity of juiciness and/or tenderness. Univariate analysis of independent attributes provides for the interpretation of treatment effects. Key attributes that differ by treatment provide researchers the ability to determine if treatments differ or not. Whereas multivariate analyses provide the opportunity for understanding how multiple variables are impacted by treatments and how consumer and trained sensory data might be related. These can be very powerful tools in understanding whether treatments affect the sensory properties of meat samples. Multivariate analysis tools provide additional information. If the dataset does not have relevant information, such as differences in ANOVA and variation is very low, however, multivariate techniques are not a way to gain information.

In using multivariate techniques, the number of observations is not as important as the strength of the relationship. For example, if a quantitative attribute used in trained panel analysis, such as cooked beef-flavor identity, does not have a quantitative relationship in univariate analysis, then these data will not be useful in multivariate analyses. There are useful resources that discuss multivariate analysis—Meilgaard et al. (2025) and Lawless and Heymann (2010). Many of the statistical packages have training that can also assist in understanding how to conduct these analyses. It is assumed that individuals using this guideline are familiar with these or comparable resources. A general description of common multivariate tools that can be used in sensory data, how they are used, what type of information they provide, and a general interpretation of data are discussed below. A dataset was extracted from Miller et al. (2023). These data include beef samples (16 treatments) that were selected to differ in flavor due to quality grade, cut, cooking method, and cooked internal cook temperature endpoint. These samples were evaluated by a trained descriptive flavor attribute panel using the Adhikari et al. (2011) method, by consumer sensory evaluation in 4 different cities (n = 80 per city), and for volatile flavor chemical analyses. Not all of the data or attributes were included in these analyses, but the data were used to illustrate how to use multivariate techniques to understand relationships within these data.

Principal component analysis

Principal component analysis (PCA) helps understand relationships among variables. Like correlation analysis, PCA studies the orthogonal relationships among variables in a dataset, but it also allows the visualization of how variables move in relation to one another. In PCA, orthogonal variation in the dataset is explained by combining variables that are related into a set of new variables, called principal components, which are uncorrelated and orthogonal to one another. The first principal component (PC 1) is derived to account for the greatest amount of variation; subsequent factors account for progressively less variation. The number of factors that can be derived is equal to n − 1, where n equals the number of variables in the dataset. Generally, only the first few factors explain significant portions of the variation, and, thus, only the first few are scrutinized. The percent of the variation that is accounted for by each factor is provided in the analysis. A bi-plot is generated that shows the relationships between factors, usually factors 1 and 2, and the sensory variables and treatments.

Figure 43 is a bi-plot generated from the aforementioned data (Miller et al., 2023). The objective is to understand relationships between beef-flavor descriptive attributes and the 16 treatments. The original data (n = 640) were averaged across panelists and treatment so that there were 16 lines of data used for PCA analysis. The first factor, or Principal Component 1, accounted for 47.3% of the variation, and the second factor, or Principal Component 2, accounted for 22.97% of the variation. The third component and subsequent components were not significant and were not presented. To interpret the bi-plot, attributes and treatments are plotted. If attributes are in the center of the plot, they are not contributing to variation in the 2 principal components. Attributes or treatments that are in a similar area are highly related. Attributes or treatments near the 0 on the vertical axis, but located away, either negatively or positively, from the horizontal axis, are accounting for more of the variation in Principal Component 2. Conversely, attributes or treatments near the 0 on the horizontal axis, but located away, either negatively or positively, from the vertical axis, account for more of the variation in Principal Component 1. Variables in the 4 quadrants between axes account for some of the variation, either negatively or positively, for both principal components. Variables used in the PCA should be measured using the same scale; these data were collected using a 16-point Spectrum Universal scale. To interpret the bi-plot, treatments clustered with variables are related. Therefore, Choice and Select bottom round roasts cooked in a crockpot to 80°C and top sirloin steaks grilled to 80°C had similar umami and beef identity. Liver-like flavor was related to Choice top loin steaks cooked on a George Foreman grill to an internal cook temperature of 80°C, whereas metallic flavor was most closely related to top sirloin steaks cooked to 58.3°C on either a George Foreman or grill. Other relationships are apparent but will not be discussed. Data where treatment may be represented by different products across companies could be analyzed for either trained or consumer sensory attributes to understand how these products differ in either trained descriptive sensory attributes or consumer attributes.

Figure 43.
Figure 43.

Principal component bi-plot of trained, beef descriptive flavor attributes from the beef-flavor lexicon (red dot) and 16 treatments (blue square), where TL, top loin steaks; GF, George Foreman grill; CTL, Choice top loin steaks; SBR, select bottom round roasts; CP, crockpot cooked; CBR, Choice bottom round roasts; and TS, top sirloin steaks. Adapted from Miller et al. (2023).

Partial least-squares regression

Partial least-squares (PLS) regression is a one-step iterative process that computes latent vectors to explain both independent and dependent variables. This tool can be used when more than one dependent variable is being predicted or when independent variables are measured using different scales. For example, if trained and consumer sensory responses were obtained on a set of samples, the trained sensory responses could be defined as the independent variables and the consumer responses defined as the dependent variables (variable being predicted so that a trained panel could be used to measure a product and understand how consumer responses would be affected). A bi-plot is shown in Figure 44, which was calculated from consumer and trained descriptive attribute sensory data for the same data presented for PCA analysis. Note that this bi-plot presents the relationships between consumer responses and trained descriptive sensory flavor attributes adapted from Miller et al. (2023).

Figure 44.
Figure 44.

Partial least-squares regression bi-plot for trained descriptive sensory attributes (red dot) to predict consumer sensory attributes (blue square), adapted from Miller et al. (2023).

These data assist the researcher in understanding what trained panel attributes to target for evaluation or that are most likely related to overall consumer liking. Once a relationship between consumer responses and trained panel responses is established, trained panel evaluation can be used to evaluate additional products at a lower cost. In this analysis, the experimental unit is the sample evaluated by each consumer and the trained descriptive panelist; a higher n is used to understand these relationships (n = 451 for the same data presented in Figure 44). The variation defined by the model is defined as 20.2% of X, which explains 45.5% of the variability in Y for the first 2 dimensions. The type of information obtained from PLS bi-plots examines variables that cluster. For example, from the bi-plot presented in Figure 44, it is apparent that consumer attributes of grilled flavor, flavor, and beef-flavor liking were highly related to overall liking. Additionally, liver-like and bitter were not related to overall consumer liking. Brown/roasted tended to be most closely related to overall consumer liking and sour basic taste, while closely related to the consumer liking attributes were most closely related to grilled flavor liking. The grouping of attributes provides an understanding of interrelationships within trained and consumer attributes and between these attributes. Different data can be used in PLS analyses, such as understanding chemical measures and either consumer or trained panel descriptive attributes.

Preference mapping

This statistical tool links consumer acceptance to trained sensory attributes and has become extensively used in the industry. The data are converted so that they are orthogonal. It is a powerful statistical tool to determine how trained panel results are related to consumers’ overall liking ratings. Preference mapping provides a rapid, objective, and reproducible tool for making product decisions. There are 2 main preference mapping tools: internal preference mapping and external preference mapping. Each tool will be discussed to provide information on how to use these tools and how to interpret these analyses.

Internal preference mapping (IMP) uses hedonic data and uses a variation of PCA. The data is converted so that they are orthogonal. On the bi-plot, instead of variables from trained panelists, consumer data and treatments are plotted. This method assists the researcher in understanding the relationship between consumer clusters or groups and overall liking and/or treatments. If consumers are clustered with a treatment, they have a preference for that treatment. An example of an IMP bi-plot is presented in Figure 45. Note that Principal Component 1 accounted for 88.79% of the variation, and Principal Component 2 accounted for 6.70% of the variation. These data are for the same dataset used for Figures 43 and 44, except that the data are presented where consumer response is the experimental unit. Consumer clusters show that not all consumers have the same preferences. In Figure 45, there are 4 general consumer clusters—one in each quadrant. Consumers in the upper right quadrant have a stronger preference for grilled flavor liking, whereas consumers in the left quadrant were furthest away from overall liking. The sensory professional could use these data to further understand why consumers are segmented as defined by pulling these consumers from the dataset and further examining their responses and/or other information obtained from these consumers, such as demographic data, one-on-one interview data, etc.

Figure 45.
Figure 45.

Internal preference mapping bi-plot for consumer overall liking attributes (red dot) and individual consumers (blue diamond). Adapted from Miller et al. (2023).

External preference mapping

External preference mapping (EPM) is used for product optimization and in product development. This tool correlates consumer preference data and trained sensory attributes and/or chemical data. Information from EPM assists the researcher in understanding what trained sensory attributes or chemical characteristics drive or are highly related to consumer likes and dislikes. The bi-plot is generated using PLS as the data were measured using different scales. Also note that the data-defined consumer is an experimental unit. While the same dataset was used in Figures 43 and 44, fewer data points were used for simplicity. Attributes are mapped in relation to consumers’ overall like ratings (Figure 46). These data show that bloody/serum-like was most closely related to consumer overall liking. However, pentanal, hexanal, bitter, and umami were negatively related to consumer overall liking. Consumers from 2 different quadrants could be identified, and further information on what was positively and negatively driving consumer acceptability, could be. These data provide direction for products that are higher in the positive and lower in the negative flavor attributes and help in understanding volatile chemical compounds that may be driving flavor.

Figure 46.
Figure 46.

External preference mapping bi-plot for consumer overall liking (green dot), trained descriptive flavor attributes and volatile flavor chemicals (red dot), and consumers (blue diamond). Adapted from Miller et al. (2023).

Other multivariate tools

New statistical and data analysis methods have been developed. The aforementioned analyses are not exclusive. Data clustering techniques, such as K-means, Agglomerative hierarchical clustering, and Gaussian mixture models, can be useful methods for understanding how respondents used scales or to segment their responses. Tools like Discriminant Analysis and machine learning tools provide further information for data analysis. As this guideline is not intended to provide an exclusive discussion of statistical methods, researchers are encouraged to discuss design and statistical methods with a statistician and/or sensory professional.

Summary

These guidelines provide guidance on cooking methods, sensory methods, and factors that influence sensory verdicts, mechanical texture methods, and data analyses of sensory data. Reference documents have been provided to assist scientists in utilizing resources that have been developed by sensory professionals and to increase understanding of methods and issues outlined in these guidelines. While these guidelines have been written to provide up-to-date methods for evaluation or application of methods to meat and meat products, new methods continue to be developed and should be considered based on their scientific merits.

Declaration of Competing Interest

Authors do not have any competing interests.

Acknowledgments

Thanks to Patty Beska and Adria Grayson Holter for assistance with the thermocouple wiring video. The development of this guide was supported in part by the American Meat Science Association Educational Development Council.

Author Contributions

Authors Miller, Papdopoulos, Schilling, and Wheeler were responsible for project administration, conceptualization, data curation, and writing the original draft, review, and editing; Gray, Dikeman, Claus, Owens, Schilling, King, Belk, Bass, Calkins, Miller, Shackelford, and Wasser contributed to conceptualization and writing the original draft, review, and editing.

Literature Cited

Adhikari, K., E. Chambers IV, R. Miller, L. Vazquez-Araujo, N. Bhumiratana, and C. Philip. 2011. Development of a lexicon for beef flavor in intact muscles. J. Sens. Stud. 26:413–420. doi: https://doi.org/10.1111/j.1745-459X.2011.00356.x

AMSA (American Meat Science Association). 1978. Guidelines for cookery and sensory evaluation of meat. American Meat Science Association, Chicago, IL.

AMSA (American Meat Science Association). 1983. Guidelines for sensory, physical and chemical measurements in ground beef. American Meat Science Association and National Live Stock and Meat Board, Chicago, IL.

AMSA (American Meat Science Association). 1995. Research guidelines for cookery, sensory evaluation and instrumental tenderness measurements of fresh meat. American Meat Science Association, Chicago, IL.

AMSA (2016) Research Guidelines for Cookery, Sensory Evaluation, and Instrumental Tenderness Measurements of Fresh Meat. 2nd Edition, American Meat Science Association. https://meatscience.org/docs/default-source/publications-resources/research-guide/amsa-research-guidelines-for-cookery-and-evaluation-1-02.pdf?sfvrsn=4c6b8eb3_2https://meatscience.org/docs/default-source/publications-resources/research-guide/amsa-research-guidelines-for-cookery-and-evaluation-1-02.pdf?sfvrsn=4c6b8eb3_2

ASTM Standard E4-24. 2010. Standard practice for force verification of testing machines. ASTM International, West Conshohocken, PA. https://www.astm.org/e4.html. (June 1, 2025).https://www.astm.org/e4.html

ASTM Standard E1871-17. 2017. Standard guide for serving protocol for sensory evaluation of foods and beverages. ASTM International, West Conshohocken, PA. https://www.astm.org/e1871-17.html. (Accessed June 1, 2025).https://www.astm.org/e1871-17.html

ASTM Standard E1885-18. 2018. Standard test method for sensory analysis—triangle test. ASTM International, West Conshohocken, PA. https://www.astm.org/e1885-18.html. (Accessed June 1, 2025).https://www.astm.org/e1885-18.html

ASTM Standard E1958-22. 2022. Standard guide for sensory claim substantiation. ASTM International, West Conshohocken, PA. https://www.astm.org/e1958-22.html. (June 1, 2025).https://www.astm.org/e1958-22.html

ASTM Standard E2299-13. 2021. Standard guide for sensory evaluation of products by children and minors. ASTM International, West Conshohocken, PA. https://www.astm.org/e2299-13.html. (Accessed June 1, 2025).https://www.astm.org/e2299-13.html

ASTM Standard E2610-18. 2018. Standard test method for sensory analysis—duo-trio test. ASTM International, West Conshohocken, PA. https://www.astm.org/e2610-18.html. (Accessed June 16, 2025).https://www.astm.org/e2610-18.html

ASTM Standard E3000-24. 2024. Standard guide for measuring and tracking performance of assessors on a descriptive panel. ASTM International, West Conshohocken, PA. https://www.astm.org/e3000-24.html. (Accessed July 16, 2025).https://www.astm.org/e3000-24.html

ASTM Standard E3009-24. 2024. Standard test method for sensory analysis—tetrad test. ASTM International, West Conshohocken, PA. https://www.astm.org/e3009-24.html. (Accessed July 16, 2025).https://www.astm.org/e3009-24.html

ASTM Standard E3041-17. 2017. Standard test method for selecting and using scales for sensory evaluation. ASTM International, West Conshohocken, PA. https://www.astm.org//e3041-17.html. (Accessed July 16, 2025).https://www.astm.org//e3041-17.html

ASTM Standard F2341-05. 2005. Standard practice for user requirements for livestock, meat, and poultry evaluation devices or systems. ASTM International, West Conshohocken, PA. https://www.astm.org/f2341-05.html. (Accessed September 15, 2024).https://www.astm.org/f2341-05.html

ASTM Standard F2343-06. 2006. Standard test method for livestock, meat, and poultry evaluation devices. ASTM International, West Conshohocken, PA. https://www.astm.org/f2343-06.html. (Accessed September 15, 2024).https://www.astm.org/f2343-06.html

ASTM MNL13-2ND-EB. 2020. Descriptive analysis testing for sensory evaluation, 2nd edition. R.N. Bleibaum (Ed.). West Conshohocken, PA: ASTM International. doi: https://doi.org/10.1520/MNL13-2ND-EB

ASTM MNL60. 2016. Physical requirement guidelines for sensory testing laboratories. 2nd Edition. C. Kuesten and L. Kruse (Eds.). West Conshohocken, PA: ASTM International. doi: https://doi.org/10.1520/MNL60-2ND-EB

ASTM MNL63. 2009. Just-about-right (JAR) scales: design, usage, benefits, and risks. L. Rothman & M. Parker (Eds.). West Conshohocken, PA: ASTM International. doi: https://doi.org/10.1520/MNL63-EB

ASTM MNL26-3RD-EB. 2020. Sensory testing methods, 3rd Edition. M. B. Wolf (Ed.). West Conshohocken, PA: ASTM International. https://www.astm.org/mnl26-3rd-eb.htmlhttps://www.astm.org/mnl26-3rd-eb.html

ASTM MNL90-2ND-EB. 2023. Guidelines for the selection and training of sensory panel members, 2nd Edition. L. Beck & T. Tamminen (Eds.). West Conshohocken, PA: ASTM International. doi: https://doi.org/10.1520/MNL90-2ND-EB

Baublits, R. T., J.-F. Meullenet, J. T. Sawyer, J. M. Mehaffey, and A. Saha. 2006. Pump rate and cooked temperature effects on pork loin instrumental, sensory descriptive and consumer-rated characteristics. Meat Sci. 72:741–750. doi: https://doi.org/10.1016/j.meatsci.2005.10.006

Belew, J. B., J. C. Brooks, D. R. McKenna, and J. W. Savell. 2003. Warner–Bratzler shear evaluations of 40 bovine muscles. Meat Sci. 64:507–512. doi: https://doi.org/10.1016/S0309-1740(02)00242-5

Berry, B. W., and M. E. Dikeman. 1994. AMSA cookery and sensory guidelines. Proc. Recip. Meat Conf. 47:51–52. https://meatscience.org/docs/default-source/publications-resources/rmc/1994/amsa-cookery-and-sensory-guidelines.pdf?sfvrsn=2. (Accessed May 7th, 2024).https://meatscience.org/docs/default-source/publications-resources/rmc/1994/amsa-cookery-and-sensory-guidelines.pdf?sfvrsn=2

Berry, B. W., and K. F. Leddy. 1990. Influence of steak temperature at the beginning of broiling on palatability, shear and cooking properties of beef loin steaks differing in marbling. J. Foodservice 5:287–297. doi: https://doi.org/10.1111/j.1745-4506.1990.tb00141.x

Beyer, E. S., L. K. Decker, E. G. Kidwell, A. L. McGinn, M. D. Chao, M. D. Zumbaugh, J. L. Vipham, and T. G. O’Quinn. 2024. Evaluation of fresh and frozen beef strip loins of equal aging periods for palatability traits. Meat Muscle Bio. 8:16903. doi: https://doi.org/10.22175/mmb.16903

Bohnenkamp, J. S., and B. W. Berry. 1987. Effect of sample numbers on sensory activity of a trained ground beef texture panel. J. Sensory Stud. 2:23–35. doi: https://doi.org/10.1111/j.1745-459X.1987.tb00184.x

Bourne, M. C. 1978. Texture profile analysis. Food Technol. 7:62–66.

Bouton, P. E., A. L. Ford, P. V. Harris, W. R. Shorthose, D. Ratcliff, and J. H. L. Morgan. 1978. Influence of animal age on the tenderness of beef: Muscle differences. Meat Sci. 2:301–311. doi: https://doi.org/10.1016/0309-1740(78)90031-1

Bradley, R. A. 1953. Some statistical methods in taste testing and quality evaluation. Biometrics. 9:22–38. doi: https://doi.org/10.2307/3001630

Cabral, A. R. 2019. Predicting beef flavor differing in lipid heat denaturation and Meillard reaction products. Ph.D. dissertation. Texas A&M University. College Station. https://oaktrust.library.tamu.edu/items/de2f2130-8545-4a11-bb2f-2317d3ac1039. (Accessed May 17, 2025).https://oaktrust.library.tamu.edu/items/de2f2130-8545-4a11-bb2f-2317d3ac1039

Cavitt, L. C., J.-F. C. Meullenet, R. Xiong, and C. M. Owens. 2005a. The correlation of Razor Blade Shear, Allo-Kramer Shear, Warner-Bratzler Shear, and sensory tests to changes in tenderness of broiler breast fillets. J. Muscle Foods 16:223–242. doi: https://doi.org/10.1111/j.1745-4573.2005.00001.x

Cavitt, L. C., C. M. Owens, J. F. Meullenet, R. K. Gandhapuneni, and G. W. Youm. 2005b. Rigor development and meat quality of large and small broilers and the use of Allo-Kramer shear, needle puncture, and razor blade shear to measure texture. Poultry Sci. 84:113–118. doi: https://doi.org/10.1093/ps/84.1.113

Cavitt, L. C., G. W. Youm, J. F. Meullenet, C. M. Owens, and R. Xiong. 2004. Prediction of poultry meat tenderness using razor blade shear, Allo-Kramer shear, and sarcomere length. J. Food Sci. 69:SNQ11–SNQ15. doi: https://doi.org/10.1111/j.1365-2621.2004.tb17879.x

Chambers, D. H., A.-M. A. Allison, and E. Chambers IV. 2004. Training effects on performance of descriptive panelists. J. Sensory Stud. 19:486–499. doi: https://doi.org/10.1111/j.1745-459X.2004.082402.x

Chambers IV, E., J. A. Bowers, and A. D. Dayton. 1981. Statistical designs and panel training/experience for sensory analysis. J. Food Sci. 46:1902–1906. doi: https://doi.org/10.1111/j.1365-2621.1981.tb04515.x

Chu, S. K. 2015. Development of an intact muscle pork flavor lexicon. MS Thesis. Texas A&M University, College Station. https://hdl.handle.net/1969.1/155014. (Accessed ??).https://hdl.handle.net/1969.1/155014

Church, I. J., and A. L. Parsons. 1993. Review: Sous-vide cook-chill technology. Intern. J. Food Sci. Technol. 28:575–586. doi: https://doi.org/10.1111/j.1365-2621.1993.tb01307.x

Claus, J. R. 1989. Characteristics of low-fat, high added water bologna. Ph.D. dissertation. Kansas State University, Manhattan.

Cochran, W. G., and G. M. Cox. 1957. Experimental design. 2nd edition. John Wiley and Sons, New York, NY.

Creed, P. G., and W. Reeve. 1998. Principles and applications of sous-vide processed foods. In: S. Ghazala, editor, Sous-vide and cool-chill processing of the food industry. Aspen Publishers, Inc, Gaithersburg, MD. p. 25–56.

Cross, H. R., R. Moen, and M. S. Stanfield. 1978. Training and testing of judges for sensory analysis of meat quality. Food Technol. 32:48–54. doi: https://doi.org/10.12691/jfnr-3-5-1

Cross, H. R., M. S. Stanfield, R. S. Elder, and G. C. Smith. 1979. A comparison of roasting versus broiling on the sensory characteristics of beef longissimus steaks. J. Food Sci. 44:310–310. doi: https://doi.org/10.1111/j.1365-2621.1979.tb10075.x

Daigle, S. P., M. W. Schilling, N. G. Marriott, H., Wang, W. E. Barbeau, and R. C. Williams. 2005. PSE-like turkey breast enhancement through adjunct incorporation in a chunked and formed deli roll. Meat Sci. 69:319–324. doi: https://doi.org/10.1016/j.meatsci.2004.08.001

Department of Health, Education, and Welfare. 1979. The Belmont report. Ethical principles and guidelines for the protection of human subjects of research. National Institutes of Health, Bethesda, MD.

Devore, J., and R. Peck. 2005. Statistics: the exploration and analysis of data. 5th edition. In: C. Crockett, editors, Summarizing bivariate data. Brooks/Cole–Thompson Learning, Inc., Belmont, CA. p. 190–191.

Djekic, I., J. M. Lorenzo, P. E. S. Munekata, M. Gagaoua, and I. Tomasevic. 2021. Review on characteristics of trained sensory panels in food science. J. Texture Stud. 52:501–509. doi: https://doi.org/10.1111/jtxs.12616

Farmer, L. J., J. M. McConnell, and D. J. Kilpatrick. 2000. Sensory characteristics of farmed and wild Atlantic salmon. Aquaculture. 187:105–125. doi: https://doi.org/10.1016/S0044-8486(99)00393-2

Field, R. A., L. C. Williams, V. S. Prasad, H. R. Cross, J. L. Secrist, and M. S. Brewer. 1984. An objective measurement for evaluation of bind in restructured lamb roasts. J. Texture Stud. 15:173–178. doi: https://doi.org/10.1111/j.1745-4603.1984.tb00377.x

Gardner, G. A., J. L. Nelson, H. G. Dolezal, J. B. Morgan, and K. K. Novotny. 1996. Effects of cooking method, endpoint temperature, and quality grade on tenderness and cooking traits of Holstein beef strip steaks. J. Anim. Sci. 74:6.

Garmyn, A. J. 2020. Consumer preferences and acceptance of meat products. Foods 9:708. doi: https://doi.org/10.3390/foods9060708

Gil, M., M. Rudy, R. Stanislawcz, and P. Duma-Kocan. 2022. Effect of traditional cooking and sous vide heat treatment, cold storage time and muscle on physicochemical and sensory properties of beef meat. Molecules 27:7307. doi: https://doi.org/10.3390/molecules27217307

Grayson, A. L., D. A. King, S. D. Shackelford, M. Koohmaraie, and T. L. Wheeler. 2014. Freezing and thawing or freezing, thawing, and aging effects on beef tenderness. J. Anim. Sci. 92:2735–2740. doi: https://doi.org/10.2527/jas.2014-7613

Harris, P. V., and W. R. Shorthose. 1988. Meat texture. In: R. A. Lawrie, editor, Developments in meat science4. Elsevier Applied Science Publishers, London, England. p. 245–286.

Hiner, R. L., L. L. Madsen, and O. G. Hankins. 1945. Histological characteristics, tenderness, and drip losses of beef in relation to temperature of freezing. J. Food Sci. 10:312–324. doi: https://doi.org/10.1111/j.1365-2621.1945.tb16173.x

Honikel, K. A. 1998. Reference methods for the physical assessment of physical characteristics of meat. Meat Sci. 49:447–457. doi: https://doi.org/10.1016/s0309-1740(98)00034-5

Hough, G., I. Wakeling, A. Mucci, E. Chambers IV, I. M. Gallardo, and L. R. Alves. 2006. Number of consumers necessary for sensory acceptability tests. Food Qual. Prefer. 17:522–526. doi: https://doi.org/10.1016/j.foodqual.2005.07.002

Hrdina-Dubsky, D. L. 1989. Sous-vide finds its niche. Food Engin. Intern. 14:40–42, 44, 46, 48.

Hunt, H. B., S. C. Watson, B. D. Chaves, and G. A. Sullivan. 2023. Inactivation of Salmonella in nonintact beef during low-temperature sous vide cooking. J. Food Protect. 86:100010. doi: https://doi.org/10.1016/j.jfp.2022.11.003

Huntsberger, D. V., and P. Billingsley. 1979. Elements of statistical inference. 4th edition. In: Analysis of variance. Allyn and Bacon, Boston, MA. p. 314–316.

IFT Sensory Evaluation Division. 1995. Guidelines for the preparation and review of papers reporting sensory evaluation data. J. Food Sci. 60:211.

Jeffe T. R., H. Wang, and E. Chambers IV. 2017. Determination of a lexicon for the sensory flavor attributes of smoked food products. J. Sens. Stud. 32:e12262. doi: https://doi.org/10.1111/joss.12262

Jeremiah, L. E., R. O. Ball, B. Uttaro, and L. L. Gibson. 1996. The relationship of chemical components to flavor attributes of bacon and ham. Food Res. Int. 29:457–464. doi: https://doi.org/10.1016/S0963-9969(96)00058-0

Jeremiah, L. E., and L. L. Gibson. 2003. Cooking influences on the palatability of roasts from the beef hip. Food Res. Int. 36:1–9. doi: https://doi.org/10.1016/S0963-9969(02)00093-5

Johnsen, P. B., G. V. Civille, and J. R. Vercellotti. 1987. A lexicon of pond-raised catfish flavor descriptors. J. Sens. Stud. 2:85–91. doi: https://doi.org/10.1111/j.1745-459X.1987.tb00190.x

King, D. A., T. L. Wheeler, S. D. Shackelford, and M. Koohmaraie. 2009. Comparison of palatability characteristics of beef gluteus medius and triceps brachii muscles. J. Anim. Sci. 87:275–284. doi: https://doi.org/10.2527/jas.2007-0809

Knight, T. D. 2006. Evaluation of frankfurters formulated with potassium lactate and sodium diacetate and inoculated with Listeria monocytogenes before and after irradiation treatment. Ph.D. Dissertation. Texas A&M University, College Station, TX. https://www.proquest.com/docview/304930511?pq-origsite=gscholar&fromopenview=true&sourcetype=Dissertations%20&%20Theses. (Accessed ??).https://www.proquest.com/docview/304930511?pq-origsite=gscholar&fromopenview=true&sourcetype=Dissertations%20&%20Theses

Kroll, B. J. 1990. Evaluating rating scales for sensory testing with children. Food Technol. 44:78–86.

Labbe, D., A. Rytz, and A. Hugi. 2004. Training is a critical step to obtain reliable product profiles in a real food industry context. Food Qual. Prefer. 15:341–348. doi: https://doi.org/10.1016/S0950-3293(03)00081-8

Laird, H. L., R. K. Miller, C. R. Kerth, M. C. Berto, and K. Adhikari. 2024. USA millennial and non-millennial beef consumers perception of beef, pork, and chicken. Meat Sci. 214:109516. doi: https://doi.org/10.1016/j.meatsci.2024.109516

Latoch, A., A. Głuchowski, and E. Czarniecka-Skubina. 2023. Sous-vide as an alternative method of cooking to improve the quality of meat: a review. Foods. 12:3110. doi: https://doi.org/10.3390/foods12163110

Lawless, H. T., and H. Heymann. 2010. Sensory evaluation of food: principles and practices. 2nd edition. Springer Science + Business Media, New York, NY.

Lawrence, T. E., D. A King, E. Obuz, E. J. Yancey, and M. E Dikeman. 2001. Evaluation of electric belt grill, forced-air convection oven, and electric broiler cookery methods for beef tenderness research. Meat Sci. 58:239–246. doi: https://doi.org/10.1016/S0309-1740(00)00159-5

Lazo, O., L. Guerrero, N. Alexi, K. Grigorakis, A. Claret, J. A. Pérez, and R. Bou. 2017. Sensory characterization, physico-chemical properties and somatic yields of five emerging fish species. Food Res. Int. 100:396–406. doi: https://doi.org/10.1016/j.foodres.2017.07.023

Lee, Y. S., C. M. Owens, and J. F. Meullenet. 2008. The Meullenet-Owens Razor Shear (MORS) for predicting poultry meat tenderness: Its applications and optimization. J. Texture Stud. 39:655–672. doi: https://doi.org/10.1111/j.1745-4603.2008.00165.x

Lowe, B., E. Crain, G. Amick, M. Riedesel, L. J. Peet, F. B. Smith, B. R. McClung, and P. S. Shearer. 1952. Defrosting and cooking frozen meat: the effect of method of defrosting and of the manner and temperature of cooking upon weight loss and palatability. In: Iowa Agricultural Experiment Station Research Bulletin. No. 385. Iowa State College, Ames, IA.

Luchak, G. L., R. K. Miller, K. E. Belk, D. S. Hale, S. A. Michaelsen, D. D. Johnson, R. L. West, F. W. Leak, H. R. Cross, and J. W. Savell. 1998. Determination of sensory, chemical and cooking characteristics of retail beef cuts differing in intramuscular and external fat. Meat Sci. 50:55–72. doi: https://doi.org/10.1016/S0309-1740(98)00016-3

Luckett, C. R., V. A. Kuttappan, L. G. Johnson, C. M. Owen, and H. S. Seo. 2014. Comparison of three instrumental methods for predicting sensory texture attributes of poultry deli meat. J. Sens. Stud. 29:171–181. doi: https://doi.org/10.1111/joss.12092

Lyon, B. G., and C. E. Lyon. 1991. Research note: shear value ranges by Instron Warner-Bratzler and single blade Allo-Kramer devices that correspond to sensory perception. Poult. Sci. 70:188–191. doi: https://doi.org/10.3382/ps.0700188

Lyon, C. E., B. G. Lyon, C. E. Davis, and W. E. Townsend. 1980. Texture profile analysis of patties made from mixed and flake-cut mechanically deboned poultry meat. Poultry Sci. 50:69–76. doi: https://doi.org/10.3382/ps.0590069

Martinez, H. A., R. K. Miller, C. R. Kerth, and B. E. Wasser. 2023. Prediction of beef tenderness and juiciness using consumer and descriptive sensory attributes. Meat Sci. 205:109292. doi: https://doi.org/10.1016/j.meatsci.2023.109292

McDowell, M. D., D. L. Harrison, C. Davey, and M. B. Stone. 1982. Differences between conventionally cooked top round roasts and semimembranosus muscle strips cooked in a model system. J. Food Sci. 47:1603–1607. doi: https://doi.org/10.1111/j.1365-2621.1982.tb04992.x

McKeith, F. M., D. D. Vol, R. Miles, P. J. Bechtel, and T. R. Carr. 1985. Chemical and sensory properties of thirteen major beef muscles. J. Food Sci. 50:869–872. doi: https://doi.org/10.1111/j.1365-2621.1985.tb12968.x

Meat Institute. 2025. Meat buyers guide. 9th edition. North American Meat Institute and American Meat Science Association, Washington, DC.

Meilgaard, M. C., G. V. Civille, and B. T. Carr. 2016. Sensory evaluation techniques. 5th edition. CRC Press, Boca Raton, FL.

Meilgaard, M. C., G. V. Civille, and B. T. Carr. 2025. Sensory evaluation techniques. 6th edition. CRC Press, Boca Raton, FL.

Miller, R. K. 2013. Personal communication. Meat Science Section, Department of Animal Science, Texas A&M University, College Station, TX.

Miller R. K., H. L. Laird, C. R. Kerth, and E. Chambers IV. 2019. The flavor and texture attributes of ground beef. Meat Muscle Biol. 1:7. doi: https://doi.org/10.221751/rmc2017.007

Miller, R. K., Gray, R. A., Kerth, C. R., Adhikari, K. & Neal, J., (2023) ‘Descriptive Beef Flavor Attributes and Consumer Acceptance Relationships for Heavy Beef Eaters”, Meat and Muscle Biology 7(1): 14449, 1–15. doi: doi: https://doi.org/10.22175/mmb.14449

Montejano, J. G., D. D. Hamann, and T. C. Lanier. 1985. Comparison of two instrumental methods with sensory texture of protein gels. J. Texture Stud. 16:403–424. doi: https://doi.org/10.1111/j.1745-4603.1985.tb00705.x

Munoz, A. M., and G. V. Civille. 1998. Universal, product and attribute scaling and the development of common lexicons in descriptive analysis. J. Sens. Stud. 13:57–75. doi: https://doi.org/10.1111/j.1745-459X.1998.tb00075.x

Munoz, A. M., G. V. Civille, and B. T. Carr. 1992. Sensory evaluation in quality control. Van Nostrand Reinhold, New York, NY.

Neely, T. R., C. L. Lorenzen, R. K. Miller, J. D. Tatum, J. W. Wise, J. F. Taylor, M. J. Buyck, J. O. Reagan, and J. W. Savell. 1999. Beef customer satisfaction: cooking method and degree of doneness effects on the top round steak. J. Anim. Sci. 77:653–660. doi: https://doi.org/10.2527/1999.773653x

Nuñez de Gonzalez, M. T., J. T. Keeton, and L. J. Ringer. 2004. Sensory and physicochemical characteristics of frankfurters containing lactate with antimicrobial surface treatments. J. Food Sci. 69:S221–S228. doi: https://doi.org/10.1111/j.1365-2621.2004.tb11009.x

O’Mahony, M. 1986. Sensory evaluation of food: Statistical methods and procedures. Marcel Dekker, Inc., New York, NY.

Ott, R. L., and M. Longnecker. 2001. An introduction to statistical methods and data analysis. 5th edition. In: C. Crockett, editor, Inferences about more than two population central values. Duxbury, Thompson Learning, Inc., Pacific Grove, CA.

Owens, C. M., and J.-F. C. Meullenet. 2010. Poultry meat tenderness. In: Handbook of poultry science and technology. John Wiley & Sons, Inc., Hoboken, New Jersey. p. 491–514. https://onlinelibrary.wiley.com/doi/pdf/10.1002/9780470504451. (Accessed ??).https://onlinelibrary.wiley.com/doi/pdf/10.1002/9780470504451

Qian, S., X. Li, C. Zhang, and C. Blecker. 2022. Effects of initial freezing rate on the changes in quality, myofibrillar protein characteristics and myowater status of beef steak during subsequent frozen storage. Int. J. Refrig. 143:148–156.

Ramsbottom, J. M., E. J. Strandine, and C. H. Koonz. 1945. Comparative tenderness of representative beef muscles. Food Res. 10:497–509. doi: https://doi.org/10.1111/j.1365-2621.1945.tb16198.x

Rhee, M. S., T. L. Wheeler, S. D. Shackelford, and M. Koohmaraie. 2004. Variation in palatability and biochemical traits within and among eleven beef muscles. J. Anim. Sci. 82:534–550. doi: https://doi.org/10.2527/2004.822534x

Sams, A. R. 1990. Electrical stimulation and high temperature conditioning of broiler carcasses. Poultry Sci. 69:1781–1786. doi: https://doi.org/10.3382/ps.0691781

Schilling, M. W., J. K. Schilling, J. R. Claus, N. G. Marriott, S. E. Duncan, and H. Wang. 2003. Instrumental texture assessment and consumer acceptability of cooked broiler breasts evaluated using a geometrically uniform-shaped sample. J. Muscle Foods. 14:11–23. doi: https://doi.org/10.1111/j.1745-4573.2003.tb00342.x

Seman, D. L., D. D. Boler, C. C. Carr, M. E. Dikeman, C. M. Owens, J. T. Keeton, T. D. Pringle, J. J. Sindelar, D. R. Woerner, A. S. de Mello, and T. H. Powell. 2018. Meat science lexicon. Meat Muscle Biol. 2. doi: https://doi.org/10.22175/mmb2017.12.0059

Shackelford, S. D., J. B. Morgan, H. R. Cross, and J. W. Savell. 1991. Identification of threshold levels for Warner-Bratzler shear force in beef top loin steaks. J. Muscle Foods. 2:289–296. doi: https://doi.org/10.1111/j.1745-4573.1991.tb00461.x

Shackelford, S. D., T. L. Wheeler, and M. Koohmaraie. 1995. Relationship between shear force and trained sensory panel tenderness ratings of 10 major muscles from Bos indicus and Bos taurus cattle. J. Anim. Sci. 73:3333–3340. doi: https://doi.org/10.2527/1995.73113333x.

Shackelford, S. D., T. L. Wheeler, and M. Koohmaraie. 1999a. Tenderness classification of beef: II. Design and analysis of a system to measure beef longissimus shear force under commercial processing conditions. J. Anim. Sci. 77:1474–1481. doi: https://doi.org/10.2527/1999.7761474x

Shackelford, S. D., T. L. Wheeler, and M. Koohmaraie. 1999b. Evaluation of slice shear force as an objective method of assessing beef longissimus tenderness. J. Anim. Sci. 77:2693–2699. doi: https://doi.org/10.2527/1999.77102693x

Shackelford, S. D., T. L. Wheeler, and M. Koohmaraie. 2004a. Technical note: Use of belt grill cookery and slice shear force for assessment of pork longissimus tenderness. J. Anim. Sci. 82:238–241. doi: https://doi.org/10.2527/2004.821238x

Shackelford, S. D., T. L. Wheeler, and M. Koohmaraie. 2004b. Evaluation of sampling, cookery, and shear force protocols for objective evaluation of lamb longissimus tenderness. J. Anim. Sci. 82:802–807. doi: https://doi.org/10.2527/2004.823802x

Sipos, L., A. Nyitrai, D. Szabo, A. Urbin, and B. V. Nagy. 2021. Former and potential developments in sensory color masking—review. Trends Food Sci. Technol. 11:1–11. doi: https://doi.org/10.1016/j.tifs.2021.02.050

Srisuradetchai, P. 2012. The analysis of square lattice designs using R and SAS. M.S. Thesis. Montana State University, Bozman, MT. https://math.montana.edu/grad_students/writing-projects/2012/12patchank.pdf. (Accessed ??).https://math.montana.edu/grad_students/writing-projects/2012/12patchank.pdf

Steel, R. G. D., and J. H. Torrie. 1980. Principles and procedures of statistics: a biometrical approach. 2nd edition. In: C. Napier and J. W. Maisel, editors, Analysis of variance I: The one-way classification. McGraw-Hill Book Company, New York, NY.

Stone, H., and J. L. Sidel. 2004. Sensory evaluation practices. 3rd edition. Elsevier Academic Press, San Diego, CA.

Tedeschi, L. O., D. Tulpan, H. M. Menendez III, and R. A. Vieira. 2025. Modeling complex systems in animal science: a pedagogical framework for new researchers. Revista Brasileira de Zootecnia 54:e20240170. doi: https://doi.org/10.37496/rbz5420240170

USDA (US Department of Agriculture). 2014. IMPS 100 fresh beef. USDA Agricultural Marketing Service, Washington, DC. http://www.ams.usda.gov/AMSv1.0/getfile?dDocName=STELDEV3003281. (Accessed ??).http://www.ams.usda.gov/AMSv1.0/getfile?dDocName=STELDEV3003281

USDA (US Department of Agriculture). 2025. Safe minimum internal temperature chart. USDA Food Safety and Inspection Service, Washington, DC. https://www.fsis.usda.gov/food-safety/safe-food-handling-and-preparation/food-safety-basics/safe-temperature-chart. (Accessed December 3, 2025).https://www.fsis.usda.gov/food-safety/safe-food-handling-and-preparation/food-safety-basics/safe-temperature-chart

USDA (US Department of Agriculture) Livestock and Poultry Program. 2012. Laboratory proficiency testing for shear force measurements. USDA, Washington, DC. https://www.ams.usda.gov/sites/default/files/media/QAD1003B_LabProficiencyTesting.pdf. (Accessed ??).https://www.ams.usda.gov/sites/default/files/media/QAD1003B_LabProficiencyTesting.pdf

Vaudagna, S. R., G. Sánchez, M. S. Neira, E. M. Insani, A. B. Picallo, M. M. Gallinger, and J. A. Lasta. 2002. Sous-vide cooked beef muscles: effects of low temperature–long time (LT–LT) treatments on their quality characteristics and storage stability. Intern. J. Food Sci. Technol. 37:425–441. doi: https://doi.org/10.1046/j.1365-2621.2002.00581.x

Wall, K. R., C. R. Kerth, R. K. Miller, and J. A. Boles. 2024. Sensory and volatile aromatic compound differences of paired lamb loins with 0 or 14 day dry aging. Small Ruminant Res. 233:107237. doi: https://doi.org/10.1016/j.smallrumres.2024.107237

Waters, C. M. 2017. Sorghum bran as an antioxidant in frozen meat and poultry products. PhD dissertation. Texas A&M University, College Station, TX. https://core.ac.uk/download/pdf/188040847.pdf. (Accessed ??).https://core.ac.uk/download/pdf/188040847.pdf

Wheeler, T. L., M. Koohmaraie, L. V. Cundiff, and M. E. Dikeman. 1994. Effects of cooking and shearing methodology on variation in Warner-Bratzler shear force values in beef. J. Anim. Sci. 72:2325–2330. doi: https://doi.org/10.2527/1994.7292325x

Wheeler, T. L., R. K. Miller, J. W. Savell, and H. R. Cross. 1990. Palatability of chilled and frozen beef steaks. J. Food Sci. 55:301–304. doi: https://doi.org/10.2527/1996.7471553x

Wheeler, T. L., S. D. Shackelford, L. P. Johnson, M. F. Miller, R. K. Miller, and M. Koohmaraie. 1997a. A comparison of Warner-Bratzler shear force assessment within and among institutions. J. Anim. Sci. 75:2423–2432. doi: https://doi.org/10.2527/1994.7292325x

Wheeler, T. L., S. D. Shackelford, and M. Koohmaraie. 1996. Sampling, cooking, and coring effects on Warner-Bratzler shear force values in beef. J. Anim. Sci. 74:1553–1562. doi: https://doi.org/10.2527/1996.7471553x

Wheeler, T. L., S. D. Shackelford, and M. Koohmaraie. 1997b. Standardizing collection and interpretation of Warner-Bratzler shear force and tenderness sensory data. Proc. Recip. Meat Conf. 50:68–77. https://www.ars.usda.gov/ARSUserFiles/30400510/1997500068.pdf. (Accessed ??).https://www.ars.usda.gov/ARSUserFiles/30400510/1997500068.pdf

Wheeler, T. L., S. D. Shackelford, and M. Koohmaraie. 1998. Cooking and palatability traits of beef longissimus steaks cooked with a belt grill or an open hearth electric broiler. J. Anim. Sci. 76:2805–2810. doi: https://doi.org/10.2527/1998.76112805x

Wheeler, T. L., S. D. Shackelford, and M. Koohmaraie. 2007. Beef longissimus slice shear force measurement among steak locations and institutions. J. Anim. Sci. 85:2283–2289. doi: https://doi.org/10.2527/jas.2006-736

Williams, E. J. 1949. Experimental designs balanced for the estimation of residual effects of treatments. Aust. J. Sci. Res., Ser. A. 2:149–168. doi: https://doi.org/10.1071/CH9490149

Zhang, Y., Y. Gao, Z. Li, Z. Zheng, X. Xu, P. Wang, B. Zheng, and Z. Qi. 2021. Correlation between instrumental stress and oral processing property of chicken broiler breast under wooden breast myopathy. Int. J. Food Sci. Technol. 56:5518–5532. doi: https://doi.org/10.1111/ijfs.15141