Are ETFs A Better Benchmark?

Jocelyn Gilligan, CFA, CIPM
Partner
June 28, 2024
15 min
Are ETFs A Better Benchmark?

Using Exchange-Traded Funds (ETFs) as benchmarks instead of traditional indices has become a common practice among investors and fund managers. ETFs offer practical advantages, such as reflecting real-world trading costs, and incorporating management fees and tax considerations. These aspects make ETFs a more accurate and accessible benchmark as they are an actual investible alternative to the strategy being assessed.

However, this approach is not without its drawbacks. Understanding both the advantages and disadvantages of using ETFs as benchmarks is crucial for making informed investment decisions and ensuring accurate performance comparisons.

This article discusses the pros and cons of using an ETF as a benchmark and considerations for making an informed decision on how to go about selecting one that is meaningful.

The Advantages:

Using an ETF as a benchmark rather than the underlying index has several advantages. These include:

Cost:

The decision to use an ETF rather than an actual index as a benchmark often stems from the costs associated with using index performance data. While index providers typically charge licensing fees for access to their indices, these fees can be cost-prohibitive for some firms, especially smaller ones, or those with limited resources.

ETFs offer a more accessible and cost-effective alternative, as they provide readily available, real-time performance data and can be traded easily on stock exchanges and accessed by anyone. By using an ETF as a benchmark, firms can circumvent the barriers to entry associated with marketing index performance directly, allowing them to still compare performance against a relevant benchmark.

Practical Investment Comparison:

ETFs represent actual investment vehicles that investors can buy and sell, thus providing a more practical and realistic performance comparison. Indices, on the other hand, are theoretical constructs that do not account for real-world trading costs, whereas ETFs do. Additionally, ETFs are traded on stock exchanges and can be bought and sold throughout the trading day at market prices, unlike indices which cannot be directly traded.

Incorporation of Costs:

ETFs include trading and management expenses and other costs associated with managing the pool of securities. When using an ETF as a benchmark, you get a more accurate reflection of the net returns an investor would actually receive after these costs. In addition, ETF performance considers the costs of buying and selling the underlying assets, including bid-ask spreads and any market impact, which indices do not.

Dividend Reinvestment:

ETFs may account for the reinvestment of dividends, providing a more accurate measure of total return. Indices often do not factor in the practical aspects of dividend reinvestment, such as timing delays, transaction costs, and tax implications, leading to a potentially less realistic depiction of investment returns.

Tax Considerations:

ETFs may have different tax treatments and efficiencies compared to the theoretical index performance. Using an ETF as a benchmark will reflect these considerations, providing a potentially more relevant comparison for taxable investors.

Replication and Tracking Error:

ETFs can exhibit tracking error, which is the deviation of the ETF's performance from the index it seeks to replicate. While tracking error may be perceived as a limitation, it also reflects the real-world challenges and frictions involved in managing an investment portfolio. Thus, using an ETF as a benchmark encompasses this aspect of real-world performance—which acknowledges the practical complexities of investing and serves to enhance transparency and accountability in investment decision making.

Transparency and Real-time Data:

ETFs provide real-time pricing information throughout trading hours, allowing investors to monitor and compare performance continuously as market conditions fluctuate. This real-time data enables more informed and timely decision-making, as investors can react instantly to market events, manage risks more effectively, and capitalize on opportunities as they arise.

Advantages Summary

In summary, using an ETF as a benchmark provides a less-costly, more realistic, practical, and accurate measure of investment performance that includes real-world considerations like costs, liquidity, tax implications, and dividend reinvestment, which are not fully captured by indices. ETFs are a true investable alternative, while indexes are not directly investible.

The Disadvantages:

While using an ETF as a benchmark has several advantages, there are also some potential drawbacks to consider:

Downside of Tracking Error:

ETFs may not perfectly track their underlying indices due to various factors such as imperfect replication methods, sampling techniques, and management decisions. This tracking error can result from differences in timing, costs, and portfolio composition between the ETF and its benchmark index.

This deviation can lead to discrepancies when comparing the ETF's performance to the actual index and can affect investors' expectations, portfolio management decisions, and performance evaluations. Thus, it is prudent to evaluate and monitor tracking error of ETFs when they are used as a benchmark.

Tracking Method: Full Replication vs. Sampling

ETFs employ different replication strategies to track their underlying indices, with some opting for full replication, while others utilize sampling techniques. These differences can lead to varying levels of tracking error and performance differences from the underlying index.

Full replication involves holding all of the securities in the index in the same proportions as they are weighted in the index, aiming to closely mirror its performance. In contrast, sampling techniques involve holding a representative subset of securities that capture the overall characteristics of the index.

While full replication theoretically offers the closest tracking to the index, it can be more costly and logistically challenging, especially for indices with a large number of securities. Sampling, while potentially more cost-effective and manageable, introduces the risk of tracking error, as the subset of securities may not perfectly reflect the index's performance.

Non-Comparable Expense Ratios:

ETFs incur management fees, which reduce returns over time. While these fees are part of the real-world costs, they can make the ETF's performance look worse compared to the theoretical performance of the index, especially when compounded over time. This may be problematic when using an ETF as a comparison tool (think expense ratios dragging down ETF benchmark performance thus making the strategy appear to have performed better than it would have against the actual index). This has the potential to influence investment decisions and performance evaluations. To address this concern, the GIPS Standards now require firms that use an ETF as a benchmark to disclose the ETF’s expense ratio.

Many active managers might argue that it’s “unfair” that the SEC requires them to compare net returns against an index that has no fees or expenses. However, if the strategy’s goal is to beat the index with active management, the manager should be doing this even after fees, otherwise passive investing (with lower fees) is a better option.

Liquidity Constraints:

Some ETFs may suffer from lower liquidity, leading to wider bid-ask spreads and higher trading costs, especially for large transactions. This can affect the ETF's performance and make it less ideal as a benchmark.

Selection Dilemma

Multiple ETFs may track the same index, each with different structures, expense ratios, and tracking accuracy (e.g., check out the differences between SPY, IVV, VOO, SPLG). As a result, choosing the most appropriate ETF as a benchmark should involve consideration of factors such as cost-effectiveness, liquidity, tracking error, and the strategy’s specific investment objectives. As a result, some due diligence should be done to ensure that the selected ETF aligns closely with the desired index and makes sense for the investment strategy.

Some firms have made it a habit to mix the use of different ETFs in factsheets, often because their data sources lack all the data needed for one ETF. While it may seem like it’s all the same, for many of the reasons discussed in this post, not all ETFs are created equal. We do not recommend mixing benchmarks, even when using actual indices (e.g., comparing performance returns to the Russell 1000 Growth, but then showing other statistics like sectors compared to the S&P 500). Similarly, we wouldn’t recommend doing that with ETFs either (e.g., comparing performance returns to IVV but using sector information from SPY). Mixing benchmark information in factsheets is messy and likely to be questioned by regulators, especially when doing so makes strategy performance look better.

Regulatory and Structural Issues:

ETFs are subject to evolving regulatory oversight that might affect their operations, costs and performance as benchmarks. This is not the case for indices.

In addition, the structural differences between ETFs, particularly regarding whether they are physically backed or use synthetic replication through derivatives, can significantly impact their risk profile and performance relative to their underlying indices.

Physically backed ETFs typically hold the actual securities that comprise the index they track, aiming to replicate its performance as closely as possible. In contrast, synthetic ETFs use derivatives, such as swaps, to replicate the index's returns without owning the underlying assets directly. While synthetic replication can offer cost and operational advantages, it also introduces counterparty risk, as the ETF relies on the financial stability of the swap provider.

As a result, it’s best to consider the structure of the ETF before using it as a benchmark.

Market Influences:

ETFs can trade at prices above (premium) or below (discount) their net asset value (NAV), which can introduce short-term performance differences that are not reflective of the underlying index performance.

These premiums and discounts arise due to supply and demand dynamics in the market, as well as factors such as investor sentiment, liquidity, and trading volume. These fluctuations can affect the ETF's reported returns and introduce discrepancies when comparing its performance to the benchmark index. Therefore, investors must consider the impact of these premiums and discounts on the ETF's short-term performance and recognize that these variances may not accurately represent the true performance of the underlying index.

When material differences in price vs. NAV exist, some firms believe that the NAV is a better representation of the fair value rather than the price and have used NAV for performance calculations. Please note that when this is done, it is important to document how fair value is determined and if the performance is based on the change in NAV or change in trading price.

Currency Risk:

Investors utilizing ETFs tracking international indices face the added complexity of currency fluctuations, which can significantly influence the ETF's performance. When investing in foreign ETFs, investors are exposed to currency risk, as fluctuations in exchange rates between the ETF's base currency and the currencies of the underlying index's constituents can impact returns. Currency movements can either enhance or detract from the ETF's performance, depending on whether the base currency strengthens or weakens relative to the underlying currencies.

Consequently, currency risk should be considered when using international ETFs as benchmarks.

Dividend Handling:

The handling of dividends by ETFs, whether they are paid out to investors or reinvested back into the fund, can have a notable impact on their total return compared to the index they track. Indices typically assume continuous reinvestment of dividends without considering real-world frictions such as transaction costs or timing delays associated with reinvestment. In contrast, ETFs may adopt different dividend distribution policies based on investor preferences and fund objectives.

ETFs that reinvest dividends back into the fund can potentially enhance their total return over time by capitalizing on the power of compounding. However, this approach may result in tracking errors if the reinvestment process incurs costs or timing discrepancies that deviate from the index's assumed reinvestment.

ETFs that distribute dividends to investors as cash payments may offer more immediate income but could lag behind the index's total return if investors do not reinvest these dividends efficiently. Therefore, the dividend handling policy adopted by an ETF can significantly influence its performance relative to the index and should be carefully considered.

Lack of Historical Data:

Some ETFs, especially newer ones, may not have a long track record. This can make historical performance comparisons less reliable or comprehensive. Without an extensive performance history, sufficient data may be lacking to assess an ETF's performance across various market conditions and economic cycles, making it challenging to gauge its potential risks and returns accurately.

Strategies that existed long before an ETF was created to track the comparable index, may end up with timing differences. Many firms often need to use multiple benchmarks to cover the entire period. But, for some strategies that go way back, an ETF may not exist back to inception. Be sure to include rationale in your documentation for benchmark selection so that it is clear when and why a benchmark was selected for the given time periods.

Conclusion:

In conclusion, using ETFs as benchmarks offers practical benefits, potentially making them a more accurate and accessible measure of investment performance compared to traditional indices since they are an actual investable alternative to hiring an active manager. However, these benefits do not come without shortcomings. By carefully evaluating these factors and considering the specifics of the ETFs selected for each strategy, managers can effectively use ETFs as benchmarks to assess and monitor investment strategies. In understanding these factors, an ETF may actually be a better comparison tool for your strategy than the underlying index.

Recommended Post

View All Articles

Most managers assume that losing an allocation comes down to returns. Underperform the benchmark, underperform peers, and the mandate goes elsewhere. That happens, but it's not usually the reason a manager gets cut from a search after the numbers already looked competitive.

More often, it's something in how the performance was presented that made an allocator hesitate. A number that didn't match across two documents. A risk statistic nobody could explain. A question in due diligence that the manager couldn't answer cleanly. None of these are calculation errors. They're trust problems, and trust is what allocators are ultimately seeking when they write a check.

Here are the performance problems we see that cost managers allocations most often, and none of them start with the returns themselves.

The Numbers Don't Match Across Documents

An allocator pulls up your factsheet, your pitchbook, and your GIPS® Composite Report, and the composite's five-year return isn't quite the same in all three. Maybe it's a rounding difference, or the factsheet reflects a different "as of" date. The allocator doesn't know that, and they aren't going to assume the best. Inconsistency reads as carelessness, and carelessness in performance reporting raises an obvious question: what else isn't being checked?

This is why we push firms to treat marketing and GIPS compliance as one coordinated process rather than two departments working from different source files. Every document that leaves the building should trace back to the same underlying data.

This matters even more now that due diligence itself is being automated. Operational due diligence teams and consultants are increasingly running AI tools that cross-check pitchbooks, factsheets, DDQs, and regulatory filings against each other, flagging contradictions that used to slip through manual review. A rounding difference or a stale figure that a person might have missed a few years ago is exactly the kind of inconsistency these tools are built to catch instantly. Clean, consistent marketing materials aren't just good practice anymore — they're what it takes to pass a review that may happen before a person ever looks at your numbers.

Performance That Looks Selected, Not Reported

Showing your best-performing account, your best-performing period, or a composite with an unusually small number of accounts invites the question every allocator is trained to ask: what am I not being shown? Due diligence teams know that everyone can't be top quartile. The SEC Marketing Rule's anti-cherry-picking provisions exist because this pattern is common enough that regulators built rules around it, and sophisticated allocators are watching for it. If your performance can be read as overly flattering rather than representative, assume a diligence team will read it that way.

Wanting to lead with your best numbers is an understandable impulse. But diligence teams are trained specifically to spot it, and selective disclosure, even when every number in it is accurate, tends to read as a bigger warning sign than an honest, complete track record would. The stronger story is discipline: the periods where you held to your stated mandate and didn't deviate even while returns lagged. That's a harder story to tell than "we outperformed," but it's the one that actually holds up, because it shows you didn't drift toward whatever was working elsewhere just to keep pace. Chasing returns outside your stated process isn't skill, it's strategy drift, and allocators are trained to spot that just as readily as cherry-picked out performance.

Our advice: resist the instinct to lead with your best examples, and show the scenarios that build trust instead. We recommend showing the ones that demonstrate you stuck to your stated mandate, policies, and procedures, especially when the outcome wasn't your best quarter. Discipline under pressure is a more durable credential than a strong one-off time period, and it's the kind of evidence that holds up long after that number is forgotten.

Statistics You Show But Can't Explain

A page full of risk statistics doesn't build confidence on its own. It invites a follow-up question, and if the manager can't explain what a downside capture ratio of 85% says about the decisions actually made in the portfolio, the statistic becomes a liability instead of an asset. Allocators aren't just checking whether the numbers are favorable. They're checking whether the manager understands their own portfolio well enough to explain it. Statistics presented without interpretation signal that the second answer is “no.”

Likewise, a page of portfolio characteristics that have nothing to do with how the strategy is actually run are not doing you any favors. If you're not making decisions at the sector level, a sector breakdown doesn't tell an allocator anything about your process. If you don't manage individual position sizing, a top-ten holdings list is not adding value.

Your factsheet should be a roadmap for the conversation you want to have, not a checklist of everything other managers include. Every number on it should be something you can explain: how it got there, what decision it reflects, and what it says about how you manage money. A statistic that's only there because everyone else shows it likely isn't helping you if it doesn’t demonstrate active decision making. It's inviting a question you may not have a good answer to.

It's the same logic as a good resume. One padded with every certification, hobby, and unrelated past role doesn't read as impressive, it reads as overwhelming and maybe irrelevant, and it makes the reader work harder to find what actually matters to the job at hand. A factsheet works the same way. The strongest ones include only what's relevant to the case being made and make it easy to connect every line back to it.

No One Can Explain Why a Decision Was Made

This is the one that costs managers the most, and it's rarely about the numbers at all. An allocator asks why a composite was redefined, why a benchmark changed, or why a particular account was excluded, and the answer is a shrug or "that's how we've always done it." Undocumented decisions create the impression that performance is being managed reactively rather than governed intentionally. Firms that can point to a clear, contemporaneous record of why a judgment call was made close that conversation quickly. Firms that can't do this will leave the allocator wondering what other judgment calls haven't been documented either.

The Common Thread

None of these problems are really about whether the strategy performed well. It comes down to whether the story behind the numbers holds up consistently under scrutiny. Allocators aren't just buying returns. They're also buying confidence that what they're being shown today will still be true, and still explainable, a year from now.

The fix isn't more disclosure for its own sake. It's making sure everything across your performance reporting tells the same, well-documented story before an allocator ever has the chance to ask why it doesn't.

GIPS® is a registered trademark owned by CFA Institute. CFA Institute does not endorse or promote this organization, nor does it warrant the accuracy or quality of the content contained herein.

There is a common assumption among boutique investment managers that the Global Investment Performance Standards (GIPS®) are built for the largest firms in the industry — that compliance is something you pursue once you've reached a certain scale, a certain client type, or a certain level of institutional credibility.

That assumption is understandable. And it is costing firms real opportunities.

The GIPS standards have no AUM threshold to get started. There is no minimum number of clients or composites required before a firm can claim compliance. And increasingly, the institutional marketplace is not waiting for firms to reach some undefined moment of readiness before asking for it. If you are newer to the GIPS standards and want a foundation for what they are and why firms pursue them, start with our post What Are the GIPS Standards?

 

The Market Has Already Decided

The gatekeepers of institutional capital such as consultants, outsourced CIO platforms, model delivery networks, and institutional allocators, have been quietly raising the bar on performance reporting standards for years. GIPS compliance has shifted from a differentiator to a baseline expectation in many of these channels.

According to eVestment, two out of three manager searches conducted by investors and consultants on their platform exclude firms that are not GIPS compliant. That means boutique managers without a compliance claim are not being passed over, they are simply not being seen. As we explored in From Compliance to Growth, GIPS compliance has effectively become the price of admission for firms seeking to expand into institutional channels.

The question is not whether your firm will eventually need it. For most managers with institutional ambitions, the answer to that question is already yes. The real question is when you choose to pursue it, and whether you make that choice on your own terms or in response to a mandate you cannot afford to lose.

 

What Compliance Actually Builds Inside Your Firm

The benefits most managers focus on are external. Things like the credibility signal, the access to channels, the due diligence box that gets checked. Those benefits are real. But some of the most meaningful returns from GIPS compliance are internal.

Implementing the GIPS standards requires firms to formalize processes that often exist informally. Composite definitions. Discretion criteria. Benchmark selection rationale. Fee policies. Error correction procedures. For many boutique managers, the implementation process is the first time these decisions have been documented and applied consistently across the firm.

That discipline matters beyond GIPS compliance itself. A firm with clean, documented performance infrastructure is better positioned for regulatory examinations, investor due diligence, and operational due diligence reviews. It demonstrates to sophisticated allocators that the firm is run with the same rigor they apply to their own oversight responsibilities. And for firms that are not primarily focused on institutional distribution, this operational foundation has standalone value, the kind of infrastructure that supports sound governance regardless of who is asking. For more on what a well-governed GIPS compliance program looks like once it is in place, see What Good GIPS Compliance Governance Looks Like in Practice.

 

The Single Best Argument for Starting Now

Here is the point that does not get made often enough: the smaller your firm and the shorter your track record, the easier it is to become compliant. That ratio flips quickly as you grow.

Retroactively constructing composites across a large number of separate accounts is genuinely difficult work, particularly when no framework existed at the time to assign accounts to composites at inception, or to move accounts between composites as investment objectives changed, client restrictions were added or removed, or mandates evolved. Working through that history portfolio by portfolio, period by period, requires both detailed documentation and sound judgment. It is one of the most time-consuming phases of any GIPS compliance implementation, and the complexity compounds with every account and every year of history added.

A firm with 30 separate accounts and a two-year track record faces a very different implementation project than the same firm a few years later with 500 accounts and a five-year track record. The strategy, the clients, and the investment process may be nearly identical, but the administrative burden of reconstructing historical composite membership correctly is not.

The firms that find implementation most manageable are the ones that started before the project grew into something unwieldy. The firms that find it most painful are the ones that waited until an institutional prospect made it urgent.

What if you are not ready to commit to full compliance yet?

That is a legitimate position. But there is a practical middle path worth considering: even if a firm does not want to claim compliance with the GIPS standards today, building out the composite structure and creating policies and procedures for managing those composites now is a worthwhile investment. That framework does not require a formal compliance claim to be useful. Additionally, it can be carried directly into a full GIPS compliance program when the time is right, dramatically reducing the effort required at that stage.

 

The Real Costs

Becoming GIPS compliant requires real work, and it is worth being direct about what that entails. At a high level, implementation comes down to four phases: defining the firm, building a GIPS standards policies and procedures manual, constructing composites and calculating performance, and creating GIPS Reports with ongoing monitoring controls. We walk through each phase in detail in A Practical Framework for Implementing the GIPS Standards.

In terms of ongoing commitment, firms should expect monthly composite management, annual GIPS Report updates, periodic policies and procedures reviews, and distribution tracking. For a lean team, owning all of this internally is often not realistic. The good news is that outsourcing to a GIPS compliance consultant is a well-established path for boutique managers and one that many firms in our client base have taken successfully. The total cost of compliance for a focused, well-organized firm is frequently lower than managers expect, particularly when implementation is approached while the firm's history and account universe are still manageable.

 

Is This the Right Time for Your Firm?

Not every firm is at the same point in this decision. Managers with the strongest case for pursuing GIPS compliance now include:

  • Firms actively pursuing institutional mandates or seeking coverage from investment consultants
  • Managers on model delivery platforms or building toward that distribution channel
  • Firms planning meaningful growth over the next two to three years
  • Any manager whose clients or prospects have already raised the question
  • Firms that simply want to build a best-in-class performance reporting foundation, regardless of where their distribution strategy stands today

The case is lower urgency for firms focused exclusively on high-net-worth or retail clients with no near-term institutional ambitions; however, there is still value in building a sound performance reporting structure, and the sooner it is established, the easier the work will be.

On Verification: You Can Wait

Verification is independent, voluntary, and valuable. It is also not required to claim compliance with the GIPS standards, and for cost-conscious boutiques, it is a reasonable place to exercise flexibility.

A firm can become GIPS compliant today and gain all the operational benefits and the ability to make the compliance claim and defer pursuing verification until there is specific demand for it. When an institutional prospect or consultant asks whether the firm is verified, that is the right moment to add it. The compliance foundation built now makes that future engagement faster and less disruptive. For a detailed walkthrough of what the verification process involves, see our series How to Survive a GIPS Verification.

Verification is worth having. It just does not need to happen on day one.

 

The Longer You Wait, The Heavier the Lift

GIPS compliance is not an initiative that gets easier with time. Every year a firm grows its account base, extends its track record, and adds complexity to its operations without a compliance framework in place is another year of history that will eventually need to be organized, documented, and reconstructed.

The managers who find implementation most straight forward are not the ones with the most resources. They are the ones who started early enough that the project was still proportionate to the size of the task.

If your firm is headed toward institutional distribution (most boutique managers we work with are), the best time to build this infrastructure is before you need it. The second best time is now.

 

Longs Peak Advisory Services specializes in GIPS compliance and investment performance consulting for investment managers and asset owners. We have helped over 250 firms implement and maintain compliance with the GIPS standards. If you are evaluating whether now is the right time for your firm, we would be glad to talk through it. Reach out athello@longspeakadvisory.com.

 

GIPS® is a registered trademark owned by CFA Institute. CFA Institute does not endorse or promote this organization, nor does it warrant the accuracy or quality of the content contained herein.

Every Spring, the performance measurement community gathers for PMAR: The Performance Measurement, Attribution & Risk Conference, hosted by TSG. This year marked the twenty-fourth annual, and I left thinking about it differently than I have in years past.

Most years, the themes evolve gradually. This year, I felt like the ground was moving.

The theme nobody put on the agenda but ran underneath nearly every session was the pace of change. Specifically, what artificial intelligence is about to do to our work. And while I came away energized, I also came away with a healthy dose of " we (as in everyone) are not ready for how fast this is coming."

Here's what stayed with me.

AI Was the Undercurrent of the Whole Event

The session titled "AI, Anxiety, and Opportunity: What Performance Professionals Need to Know" was, predictably, one of the most sought-after sessions of the conference. The panel, which included practitioners from across the industry, did a nice job naming both sides of the coin: the anxiety of not knowing what your job looks like in five years, and the opportunity sitting right in front of us if we lean in.

Here's my honest read of the room, though. The mood was optimistic. Maybe a little too optimistic. There was a comfortable assumption that AI will mostly handle the tedious parts and leave the interesting work to us. Or that AI won’t take your job, someone that knows AI will. I'm not sure it'll be that tidy.

From what we're already seeing in our own work and across the firms we serve, the capabilities are advancing faster than most people can comprehend. The days where “our industry is just slower to adapt” are gone. Just last week, anthropic released Fable 5 and before it was shut down (temporarily?), we played around with it a little and its capabilities are dumbfounding. I don't think it will be long before these conferences look drastically different. Different sessions, different vendors, maybe a different sense of what the job even is. That's not a doom prediction. It's just a reason to pay closer attention than feels comfortable.

Separating Skill From Luck Just Got Harder and More Important

One of my favorite sessions was Michael Ervolini's "You Can't Find Skill in Returns: Distinguishing Performance From the Decisions That Generate Them." It's a deceptively simple premise: returns tell you what happened, not whether the manager was actually good. A great number can come from a great decision, or from luck. A bad number can hide genuine skill.

What I appreciate about PMAR is that the community keeps bringing fresh perspectives to this old, hard problem: how do we actually evaluate skill versus luck? It's a question that never fully resolves, and every year someone pushes the thinking forward.

It struck me that this question gets more important in an AI world, not less. As machines take over more of the calculation and even some of the decision-making, our value shifts toward judgment – knowing which decisions deserved credit, which results were noise, and what a number actually means in context. That's the kind of discernment a model can assist with but can't own. For more from Mr. Ervolini, here's a link to his latest book Skill vs. Luck.

The GIPS Challenges That Keep Coming Back

I'm biased here, but the "Common GIPS Challenges and How to Avoid Them" session was a highlight for us, in part because our own Matthew Deatherage, CFA, CIPM, was on the panel alongside peers from TSG, MassPRIM, and Strategic Investment Group.

What I always find striking about this topic is how consistent the challenges are. Firms pursuing compliance with the Global Investment Performance Standards (GIPS®)* tend to stumble on the same handful of issues year after year, and almost all of them are avoidable with the right foundation in place. That's a big part of why we do what we do at Longs Peak: helping firms get ahead of those pitfalls instead of discovering them during verification or, worse, during a regulatory exam.

Matt is a familiar face on these panels, and it's great to have our perspective in the mix. But the takeaway that stuck with me tied right back to the AI thread running through the whole conference.

Across several different panels, presenters talked about feeding the GIPS standards into their own AI models to churn out GIPS reports. And here's the thing, anyone can do that. You can drop the standards into a model in minutes. What a model can't do is provide critical judgment about how a principles-based framework should be applied to your specific facts and circumstances and whether those GIPS reports and statistics were calculated correctly. The GIPS standards aren't a checklist; they're a set of principles that require interpretation, and interpretation is exactly where experience earns its keep.

I'm not saying don't use AI to help build a framework. Use it. But like any model, if you don't really know what you're asking it to do, the output won't save you. Simply asking a model to "make my firm GIPS compliant" isn't going to make it so. At least not yet!

And there's one problem every performance professional already knows AI hasn't solved: data. As they say, garbage in, garbage out. Meaningful performance lives and dies on clean, well-organized data, and no software tool or AI model fixes messy inputs alone. At Longs Peak, we have spent the last 10 years working with clients to improve data quality through data integrity testing. For us, these AI models have only expanded what’s possible. We know one thing for sure: setting these tools up with the proper context (i.e., knowing what to look for) and then evaluating that context on an ongoing basis may turn out to be the most crucial piece of it all.

CFA Institute Is Listening on the CIPM

A session I didn't expect to find as interesting as I did was "CIPM Through the Practitioner Lens," facilitated by Rob Langrick of CFA Institute. Rather than simply presenting at the room, CFA Institute came to listen and gather candid feedback on the CIPM designation: where it's delivering value, where it's falling short, and how it should evolve to stay relevant to the work we actually do day to day.

The audience didn't hold back, and there were some genuinely thoughtful suggestions including how the code of ethics will evolve in this new AI era, some recommendations on reformatting the exam to break it into smaller chunks (going into greater detail on each) as well as adding a CIPM group within the CFA societies to encourage further connection. It was refreshing to see CFA Institute putting real energy behind a credential that so many of us have invested in and want to see grow in value. Given the pace of change in our field, willingness to adapt feels necessary. For anyone interested in contributing ideas to the CIPM, you can use this link to provide feedback.

A Quick Word on the Trivia

I'd be remiss not to mention that Performance Trivia got a much-needed upgrade this year. In past years, only a handful of contestants got to play while the rest of us watched (though in fairness, not all of us were clamoring for the spotlight). The new format this time allowed everyone to participate (without taking center stage), and it was a lot more fun for it. A small change, but it captured something I value about this community: it's competitive, but it's also genuinely collegial and prides itself on memorizing quirky names and vintage formulas.

Before PMAR Even Started: Women in Performance Measurement

For me, the week actually started the day before the conference, at the Women in Performance Measurement (WiPM) gathering. An event created just for the women in our industry. It's one of my favorite parts of this trip every year, and not only because the conversation is good. There's something energizing about being in a room full of women who do this work, comparing notes and reconnecting.

Fittingly, AI came up here too, though in a much more hands-on way than it would on the main stage. Practitioners shared real use cases, both personal and professional: the small ways AI is already saving them time day to day, and the bigger experiments they're running at their firms. It was practical, curious, and refreshingly free of hype.

We were also lucky to have a guest speaker, Lidia Arshavsky, who spoke on executive presence. She broke down how executive presence actually gets evaluated inside organizations (the signals people pick up on, often without realizing it) and offered practical recommendations for strengthening your own. It was the kind of talk that's useful no matter where you are in your career.

It was a great way to kick off PMAR, and an even better way to reconnect with women I only get to see a few times a year. Sometimes the most valuable part of a conference happens in these opportunities to network and reconnect within our niche performance community. A big thank you to TSG who donated the space for this event to take place and have done so for many years.

What AI Can't Take From Us

The conference's forward-looking sessions, including "Innovative Ways to Present Performance: Dashboards & Analytics," got me thinking. The tools are evolving so quickly and so much of the analysis, presentation, and reporting can now be automated. I am left wondering how long the traditional use of software in our space will last in its current form.

When the capabilities advancing fastest don’t always come from the established vendors, who benefits? My hope is that everyone does. That these tools level a playing field that used to tilt heavily toward the largest institutions, give smaller firms the ability to deliver high-caliber analytics previously out of reach, and push the whole field toward better solutions. That makes for a more competitive space and ultimately a clearer picture for investors to evaluate their options.

That's the optimistic case, and I believe it. But it only holds if we stay clear-eyed about where our own value comes from and that's the note I want to leave you on. The pace of change is a reason to focus, not to panic. The things that make us valuable are the things AI can't take: consciousness, judgment, and the human-in-the-loop accountability that clients ultimately trust. Machines will calculate faster and present prettier. They won't sit across the table from a client and take responsibility for what a number actually means.

So, by all means, get curious about the tools (Claude seemed to be most people’s favorite – mine as well). Experiment. Don't be the individual or firm that gets left behind. But anchor yourself in the part of this work that's irreplaceably human, because that's the part that was always the point.

See you at PMAR 2027. I suspect it'll look a little different.

________________________________________

GIPS® is a registered trademark owned by CFA Institute. CFA Institute does not endorse or promote this organization, nor does it warrant the accuracy or quality of the content contained herein.